Abstract
Deep learning has emerged as a cornerstone technology across various domains, from image classification to natural language processing. However, the computational and data demands of training large-scale neural networks pose significant challenges. Distributed learning approaches, particularly those leveraging data parallelism, have become critical to addressing these challenges. Among these, the parameter server architecture stands out as a widely adopted and scalable solution, enabling efficient training of large models across distributed systems. This survey provides a comprehensive exploration of the parameter server architecture, detailing its design principles and operation. It categorizes and critically analyzes research advancements across five key aspects: consistency control, network optimization, parameter management, straggler handling, and fault tolerance. By synthesizing insights from a wide range of studies, this work highlights the trade-offs and practical effectiveness of various approaches while identifying open challenges and future research directions. The survey aims to serve as a foundational resource for researchers and practitioners striving to enhance the performance and scalability of distributed deep learning systems.
Author supplied keywords
Cite
CITATION STYLE
Provatas, N., Konstantinou, I., & Koziris, N. (2025). A Survey on Parameter Server Architecture: Approaches for Optimizing Distributed Centralized Learning. IEEE Access, 13, 30993–31015. https://doi.org/10.1109/ACCESS.2025.3535085
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.