Overview of Memory-Efficient Architectures for Deep Learning in Real-Time Systems †

3Citations
Citations of this article
9Readers
Mendeley users who have this article in their library.
Get full text

Abstract

With advancements in artificial intelligence (AI), deep learning (DL) has become crucial for real-time data analytics in areas like autonomous driving, healthcare, and predictive maintenance; however, its computational and memory demands often exceed the capabilities of low-end devices. This paper explores optimizing deep learning architectures for memory efficiency to enable real-time computation in low-power designs. Strategies include model compression, quantization, and efficient network designs. Techniques such as eliminating unnecessary parameters, sparse representations, and optimized data handling significantly enhance system performance. The design addresses cache utilization, memory hierarchies, and data movement, reducing latency and energy use. By comparing memory management methods, this study highlights dynamic pruning and adaptive compression as effective solutions for improving efficiency and performance. These findings guide the development of accurate, power-efficient deep learning systems for real-time applications, unlocking new possibilities for edge and embedded AI.

Cite

CITATION STYLE

APA

Demir, B., Domazet, E., & Mechkaroska, D. (2025). Overview of Memory-Efficient Architectures for Deep Learning in Real-Time Systems †. Engineering Proceedings, 104(1). https://doi.org/10.3390/engproc2025104077

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free