Abstract
This paper extends the reach of General Purpose GPU programming by presenting a software architecture that supports efficient fine-grained synchronization over global memory. The key idea is to transform global synchronization into global communication so that conflicts are serialized at the thread block level.With this structure, the threads within each thread block can synchronize using low latency, high-bandwidth local scratchpad memory. To enable this architecture, we implement a scalable and efficient message passing library. Using Nvidia GTX 1080 ti GPUs, we evaluate our new software architecture by using it to solve a set of five irregular problems on a variety of workloads.We find that on average, our solutions improve performance over carefully tuned state-of-the-art solutions by 3.6×.
Author supplied keywords
Cite
CITATION STYLE
Wang, K., Fussell, D., & Lin, C. (2019). Fast Fine-Grained Global Synchronization on GPUs. In International Conference on Architectural Support for Programming Languages and Operating Systems - ASPLOS (pp. 793–806). Association for Computing Machinery. https://doi.org/10.1145/3297858.3304055
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.