Abstract
Reinforcement learning (RL) offers a promising, adaptive alternative to heuristic-based autoscaling, yet its practical adoption in production environments remains negligible. In this paper, we argue that this gap between promise and practice is caused by three systemic challenges that violate fundamental RL assumptions: (i) generalization failures under workload and system drift; (ii) orchestration interference that obscures causality; and (iii) unreliable, delayed metric feedback. We substantiate these claims through an empirical study of two PPO-based autoscalers on real-world and synthetic workloads, demonstrating how these factors lead to policy instability and performance degradation. Our findings reveal that these challenges collectively frame autoscaling as a Partially Observable Markov Decision Process. We conclude that robust RL-based autoscaling requires a paradigm shift from purely algorithmic solutions toward systems-aware designs that model the partial observability and non-stationarity inherent in service autoscaling.
Author supplied keywords
Cite
CITATION STYLE
Asadi, N., Ali, D., Ursu, R. M., & Kellerer, W. (2025). Challenges in Designing Robust RL-Based Autoscalers. In PACMI 2025 - Proceedings of the 4th Workshop on Practical Adoption Challenges of ML for Systems (pp. 44–49). Association for Computing Machinery, Inc. https://doi.org/10.1145/3766882.3767176
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.