Abstract
Graph neural networks (GNN) inferencing involves weighting vertex feature vectors, followed by aggregating weighted vectors over a vertex neighborhood. High and variable sparsity in the input vertex feature vectors, and high sparsity and power-law degree distributions in the adjacency matrix, can lead to (a) unbalanced loads and (b) inefficient random memory accesses. GNNIE ensures load-balancing by splitting features into blocks, proposing a flexible MAC architecture, and employing load (re)distribution. GNNIE's novel caching scheme bypasses the high costs of random DRAM accesses. GNNIE shows high speedups over CPUs/GPUs; it is faster and runs a broader range of GNNs than existing accelerators.
Author supplied keywords
Cite
CITATION STYLE
Mondal, S., Manasi, S. D., Kunal, K., Ramprasath, S., & Sapatnekar, S. S. (2022). GNNIE: GNN Inference Engine with Load-balancing and Graph-specific Caching. In Proceedings - Design Automation Conference (pp. 565–570). Institute of Electrical and Electronics Engineers Inc. https://doi.org/10.1145/3489517.3530503
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.