Abstract
Increasingly larger and better Transformer models keep advancing state-of-the-art accuracy and capability for Natural Language Processing applications. These models demand more computational power, storage, and energy. Mokey reduces the footprint of stateof-the-art 32-bit or 16-bit floating-point transformer models by quantizing all values to 4-bit indexes into dictionaries of representative 16-bit fxed-point centroids. Mokey does not need fne-tuning, an essential feature as often the training resources or datasets are not available to many. Exploiting the range of values that naturally occur in transformer models, Mokey selects centroid values to also ft an exponential curve. This unique feature enables Mokey to replace the bulk of the original multiply-accumulate operations with narrow 3b fxed-point additions resulting in an area-and energy-efcient hardware accelerator design. Over a set of stateof-the-art transformer models, the Mokey accelerator delivers an order of magnitude improvements in energy efciency over a Tensor Cores-based accelerator while improving performance by at least 4× and as much as 15× depending on the model and on-chip buffering capacity. Optionally, Mokey can be used as memory compression assist for any other accelerator transparently stashing wide floating-point or fxed-point activations or weights into narrow 4-bit indexes. Mokey proves superior to prior state-of-the-art quantization methods for Transformers.
Author supplied keywords
Cite
CITATION STYLE
Zadeh, A. H., Mahmoud, M., Abdelhadi, A., & Moshovos, A. (2022). Mokey: Enabling Narrow Fixed-Point Inference for Out-of-the-Box Floating-Point Transformer Models. In Proceedings - International Symposium on Computer Architecture (pp. 888–901). Institute of Electrical and Electronics Engineers Inc. https://doi.org/10.1145/3470496.3527438
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.