Quantized Object Detection for Real-Time Inference on Embedded GPU Architectures

9Citations
Citations of this article
12Readers
Mendeley users who have this article in their library.

Abstract

Deploying deep learning-based object detection models like YOLOv4 on resource-constrained embedded architectures presents several challenges, particularly regarding computing performance, memory usage, and energy consumption. This study examines the quantization of the YOLOv4 model to facilitate real-time inference on lightweight edge devices, focusing on NVIDIA’s Jetson Nano and AGX. We utilize posttraining quantization techniques to reduce both model size and computational complexity, all while striving to maintain acceptable detection accuracy. Experimental results indicate that an 8-bit quantized YOLOv4 model can achieve near real-time performance with minimal accuracy loss. This makes it wellsuited for embedded applications such as autonomous navigation. Additionally, this research highlights the trade-offs between model compression and detection performance, proposing an optimization method tailored to the hardware constraints of embedded architectures.

Cite

CITATION STYLE

APA

Guerrouj, F. Z., Flórez, S. R., El Ouardi, A., Abouzahir, M., & Ramzi, M. (2025). Quantized Object Detection for Real-Time Inference on Embedded GPU Architectures. International Journal of Advanced Computer Science and Applications, 16(5), 20–29. https://doi.org/10.14569/IJACSA.2025.0160503

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free