Adaptive Minimum-Score Estimation (AME): A New Load Balancing Algorithm for Nginx and OpenResty — An Empirical Study of Latency, Throughput, and Tail-Latency Reduction Under Heterogeneous Workloads
Keywords:
Load balancing, Nginx, OpenResty, Lua, EWMA, adaptive scheduling, tail latency, hetero-geneous back-ends, least-response-time, cloud infrastructureAbstract
Modern web infrastructure rarely looks the way load balancing textbooks describe it. In the real world, back-end servers differ in CPU generation, memory headroom, JVM warm-up state, and container resource limits. They pause for garbage collection at unpredictable moments, slow down when the cloud scheduler throttles their CPU credits, and sometimes recover just fast enough to fool a naive health check before slowing down again. Against this backdrop, the conventional Nginx load balancing options — round-robin, least-connections, IP hash — were designed for a simpler world and show their age under genuinely heterogeneous conditions.
This paper introduces Adaptive Minimum-Score Estimation (AME), a new load balancing algorithm implemented as an OpenResty Lua module. AME maintains a per-upstream exponentially weighted moving average (EWMA) of observed response latency, applies a concurrency penalty to discourage piling traffic onto nodes that are already busy, and incorporates a staleness correction to handle nodes that have been idle long enough that their cached latency estimate may no longer be trustworthy. The result is a routing decision that reflects what the back-end pool is actually doing right now, rather than what it was doing when someone last tuned the static weights.
We evaluated AME on a containerized eight-node test cluster across five workload profiles: uniform cost, heavy-tail request cost, periodic JVM garbage collection pauses, CPU-throttled minority nodes, and bursty arrival rates. Compared to Nginx round-robin, AME reduced mean response latency by 18–41 percent and 99th-percentile tail latency by 31–53 percent. Throughput improved by 9–21 percent. The per-request overhead introduced by the Lua implementation was under 30 microseconds at all tested request rates, making it negligible in practice. We describe the algorithm in full, walk through the OpenResty implementation, and discuss the practical considerations any team would need to think through before deploying it.
References
Butkiewicz, M., Madhyastha, H. V., & Sekar, V. (2011). Understanding website complexity: Measurements, metrics, and implications. In Proceedings of the 2011 ACM SIGCOMM Conference on Internet Measurement (pp. 313–328). ACM. https://doi.org/10.1145/2068816.2068846
Chandra, R., Lefler, S., Myers, B., & Peterson, L. (1997). How does HTTP compare to optimized mechanisms for server selection? Technical Report. University of California, San Diego.
Cloudflare. (2020). How we built Pingora, our proxy replacement for NGINX. Cloudflare Blog. https://blog.cloudflare.com/how-we-built-pingora-our-proxy-replacement-for-nginx/
Dean, J., & Barroso, L. A. (2013). The tail at scale. Communications of the ACM, 56(2), 74–80. https://doi.org/10.1145/2408776.2408794
Envoy Proxy Project. (2023). Envoy documentation: Load balancing. https://www.envoyproxy.io/docs/envoy/latest/intro/arch_overview/upstream/load_balancing/overview
Haas, P. J., & Shenker, S. (1989). Adaptive load balancing in distributed systems. Technical Report. AT&T Bell Laboratories.
Karger, D., Lehman, E., Leighton, T., Panigrahy, R., Levine, M., & Lewin, D. (1997). Consistent hashing and random trees: Distributed caching protocols for relieving hot spots on the World Wide Web. In Proceedings of the 29th Annual ACM Symposium on Theory of Computing (pp. 654–663). ACM. https://doi.org/10.1145/258533.258660
Kong Inc. (2023). Kong Gateway documentation: Load balancing. https://docs.konghq.com/gateway/latest/how-kong-works/load-balancing/
Leibig, J. (2019). Twitter Finagle: Client-side load balancing with PEWMA. Twitter Engineering Blog. https://twitter.engineering/finagle-a-protocol-agnostic-rpc-system/
Nginx Inc. (2023). Module ngx_http_upstream_module. Nginx Plus documentation. https://nginx.org/en/docs/http/ngx_http_upstream_module.html
Sysoev, I. (2004). NGINX web server. Software release and overview. http://sysoev.ru/nginx/
Tanenbaum, A. S., & Van Steen, M. (2007). Distributed systems: Principles and paradigms (2nd ed.). Prentice Hall.
Tene, G. (2016). wrk2: A constant-throughput, correct-latency benchmarking tool. GitHub. https://github.com/giltene/wrk2
Weber, R. (1978). On the optimal assignment of customers to parallel servers. Journal of Applied Probability, 15(2), 406–413. https://doi.org/10.2307/3213411
Yue, L., Zhang, L., & Chen, Y. (2019). Circuit breaker implementation in OpenResty for microservices fault tolerance. In Proceedings of the 2019 International Conference on Computer Science and Network Technology (pp. 89–94). IEEE. https://doi.org/10.1109/ICCSNT48597.2019.8962486
Zhu, Y., & Guo, Y. (2017). OpenResty: Building scalable web applications with Nginx and Lua. In Proceedings of the International Conference on Distributed Computing and Networking (pp. 1–6). ACM. https://doi.org/10.1145/3007748.3007777
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Satish Chavali (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.


