Abstract
We consider the problem of how to improve memory latency tolerance inmassively multithreaded GPGPUs when the thread-level parallelism of an application is not sufficient to hide memory latency. One solution used in conventional CPU systems is prefetching, both in hardware and software. However, we show that straightforwardly applying such mechanisms to GPGPU systems does not deliver the expected performance benefits and can in fact hurt performance when not used judiciously. This paper proposes new hardware and software prefetching mechanisms tailored to GPGPU systems, which we refer to as many-thread aware prefetching (MT-prefetching) mechanisms. Our software MT-prefetching mechanism, called interthread prefetching, exploits the existence of common memory access behavior among fine-grained threads. For hardware MT-prefetching, we describe a scalable prefetcher training algorithm along with a hardware-based inter-thread prefetching mechanism. In some cases, blindly applying prefetching degrades performance. To reduce such negative effects, we propose an adaptive prefetch throttling scheme, which permits automatic GPGPU application- and hardware-specific adjustment. We show that adaptation reduces the negative effects of prefetching and can even improve performance. Overall, compared to the state-of-the-art software and hardware prefetching, our MT-prefetching improves performance on average by 16%(software pref.) / 15% (hardware pref.) on our benchmarks. © 2010 IEEE.
Author supplied keywords
Cite
CITATION STYLE
Lee, J., Lakshminarayana, N. B., Kim, H., & Vuduc, R. (2010). Many-thread aware prefetching mechanisms for GPGPU applications. In Proceedings of the Annual International Symposium on Microarchitecture, MICRO (pp. 213–224). https://doi.org/10.1109/MICRO.2010.44
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.