Adaptive Minimum-Score Estimation (AME): A New Load Balancing Algorithm for Nginx and OpenResty — An Empirical Study of Latency, Throughput, and Tail-Latency Reduction Under Heterogeneous Workloads

Authors

  • Satish Chavali SAP America INC, USA. Author

Keywords:

Load balancing, Nginx, OpenResty, Lua, EWMA, adaptive scheduling, tail latency, hetero-geneous back-ends, least-response-time, cloud infrastructure

Abstract

Modern web infrastructure rarely looks the way load balancing textbooks describe it. In the real world, back-end servers differ in CPU generation, memory headroom, JVM warm-up state, and container resource limits. They pause for garbage collection at unpredictable moments, slow down when the cloud scheduler throttles their CPU credits, and sometimes recover just fast enough to fool a naive health check before slowing down again. Against this backdrop, the conventional Nginx load balancing options — round-robin, least-connections, IP hash — were designed for a simpler world and show their age under genuinely heterogeneous conditions.
This paper introduces Adaptive Minimum-Score Estimation (AME), a new load balancing algorithm implemented as an OpenResty Lua module. AME maintains a per-upstream exponentially weighted moving average (EWMA) of observed response latency, applies a concurrency penalty to discourage piling traffic onto nodes that are already busy, and incorporates a staleness correction to handle nodes that have been idle long enough that their cached latency estimate may no longer be trustworthy. The result is a routing decision that reflects what the back-end pool is actually doing right now, rather than what it was doing when someone last tuned the static weights.
We evaluated AME on a containerized eight-node test cluster across five workload profiles: uniform cost, heavy-tail request cost, periodic JVM garbage collection pauses, CPU-throttled minority nodes, and bursty arrival rates. Compared to Nginx round-robin, AME reduced mean response latency by 18–41 percent and 99th-percentile tail latency by 31–53 percent. Throughput improved by 9–21 percent. The per-request overhead introduced by the Lua implementation was under 30 microseconds at all tested request rates, making it negligible in practice. We describe the algorithm in full, walk through the OpenResty implementation, and discuss the practical considerations any team would need to think through before deploying it.

References

Butkiewicz, M., Madhyastha, H. V., & Sekar, V. (2011). Understanding website complexity: Measurements, metrics, and implications. In Proceedings of the 2011 ACM SIGCOMM Conference on Internet Measurement (pp. 313–328). ACM. https://doi.org/10.1145/2068816.2068846

Chandra, R., Lefler, S., Myers, B., & Peterson, L. (1997). How does HTTP compare to optimized mechanisms for server selection? Technical Report. University of California, San Diego.

Cloudflare. (2020). How we built Pingora, our proxy replacement for NGINX. Cloudflare Blog. https://blog.cloudflare.com/how-we-built-pingora-our-proxy-replacement-for-nginx/

Dean, J., & Barroso, L. A. (2013). The tail at scale. Communications of the ACM, 56(2), 74–80. https://doi.org/10.1145/2408776.2408794

Envoy Proxy Project. (2023). Envoy documentation: Load balancing. https://www.envoyproxy.io/docs/envoy/latest/intro/arch_overview/upstream/load_balancing/overview

Haas, P. J., & Shenker, S. (1989). Adaptive load balancing in distributed systems. Technical Report. AT&T Bell Laboratories.

Karger, D., Lehman, E., Leighton, T., Panigrahy, R., Levine, M., & Lewin, D. (1997). Consistent hashing and random trees: Distributed caching protocols for relieving hot spots on the World Wide Web. In Proceedings of the 29th Annual ACM Symposium on Theory of Computing (pp. 654–663). ACM. https://doi.org/10.1145/258533.258660

Kong Inc. (2023). Kong Gateway documentation: Load balancing. https://docs.konghq.com/gateway/latest/how-kong-works/load-balancing/

Leibig, J. (2019). Twitter Finagle: Client-side load balancing with PEWMA. Twitter Engineering Blog. https://twitter.engineering/finagle-a-protocol-agnostic-rpc-system/

Nginx Inc. (2023). Module ngx_http_upstream_module. Nginx Plus documentation. https://nginx.org/en/docs/http/ngx_http_upstream_module.html

Sysoev, I. (2004). NGINX web server. Software release and overview. http://sysoev.ru/nginx/

Tanenbaum, A. S., & Van Steen, M. (2007). Distributed systems: Principles and paradigms (2nd ed.). Prentice Hall.

Tene, G. (2016). wrk2: A constant-throughput, correct-latency benchmarking tool. GitHub. https://github.com/giltene/wrk2

Weber, R. (1978). On the optimal assignment of customers to parallel servers. Journal of Applied Probability, 15(2), 406–413. https://doi.org/10.2307/3213411

Yue, L., Zhang, L., & Chen, Y. (2019). Circuit breaker implementation in OpenResty for microservices fault tolerance. In Proceedings of the 2019 International Conference on Computer Science and Network Technology (pp. 89–94). IEEE. https://doi.org/10.1109/ICCSNT48597.2019.8962486

Zhu, Y., & Guo, Y. (2017). OpenResty: Building scalable web applications with Nginx and Lua. In Proceedings of the International Conference on Distributed Computing and Networking (pp. 1–6). ACM. https://doi.org/10.1145/3007748.3007777

Downloads

Published

2026-07-06

How to Cite

Adaptive Minimum-Score Estimation (AME): A New Load Balancing Algorithm for Nginx and OpenResty — An Empirical Study of Latency, Throughput, and Tail-Latency Reduction Under Heterogeneous Workloads. (2026). ISCSITR- INTERNATIONAL JOURNAL OF COMPUTER SCIENCE AND ENGINEERING (ISCSITR-IJCSE) - ISSN: 3067-7394, 7(2), 1-35. https://iscsitr.in/index.php/ISCSITR-IJCSE/article/view/ISCSITR-IJCSE_2026_07_02_001