Abstract
Dataflow architectures are considered promising architecture, offering a commendable balance of performance, efficiency, and flexibility. Abundant prior works have been proposed to improve the performance of dataflow architectures. Nevertheless, these solutions can be further improved due to the lack of efficient data prefetching and flexible task scheduling. In this article, we propose a novel dataflow architecture with adaptive prefetching and decentralized scheduling (PANDA). First, we present an application-adaptive data prefetching method and on-chip memory microarchitecture designed to overlap memory access latency. Second, we introduce a decentralized dataflow scheduling approach and processing element (PE) microarchitecture aimed at improving hardware utilization. Experimental results show that in a wide range of real-world applications, PANDA attains up to 2.53× performance improvement and 1.79× energy efficiency improvement over the state-of-the-art dataflow architectures.
Author supplied keywords
Cite
CITATION STYLE
Qin, S., Fan, Z., Li, W., Wang, Z., An, X., Ye, X., & Fan, D. (2025). PANDA: Adaptive Prefetching and Decentralized Scheduling for Dataflow Architectures. ACM Transactions on Architecture and Code Optimization, 22(2). https://doi.org/10.1145/3721288
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.