Abstract
This article presents GRAM ( G PU-based R untime A daption for M ixed-precision) a framework for the effective use of mixed precision arithmetic for CUDA programs. Our method provides a fine-grain tradeoff between output error and performance. It can create many variants that satisfy different accuracy requirements by assigning different groups of threads to different precision levels adaptively at runtime . To widen the range of applications that can benefit from its approximation, GRAM comes with an optional half-precision approximate math library. Using GRAM, we can trade off precision for any performance improvement of up to 540%, depending on the application and accuracy requirement.
Cite
CITATION STYLE
Ho, N.-M., silva, H. D., & Wong, W.-F. (2021). GRAM. ACM Transactions on Architecture and Code Optimization, 18(2), 1–24. https://doi.org/10.1145/3441830
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.