Abstract
Recent advancements in Silicon Photonics (SiPh)-based AI accelerators offer promising solutions to address the energy bottlenecks of executing large models, such as Transformers. However, existing SiPh solutions primarily focus on accelerating matrix multiplication (MM) in the photonic domain, while the Softmax Activation (SMA) function - responsible for 30-40% of the total computation in Transformers - is still handled by conventional digital platforms. This reliance on digital platforms introduces significant energy and latency overheads due to the frequent conversions between photonic MM and digital SMA.While nonlinear functions like ReLU, Sigmoid, Tanh, and Softplus have been successfully implemented in the photonic domain using components such as optical amplifiers, Mach-Zehnder modulators (MZMs), and microring resonators (MRRs), similar approaches are not suitable for SMA. The excessive area and energy overhead introduced by optical amplifiers, along with the inability of MRRs and MZMs to handle SMA's exponential and division functions, make it difficult to realize a SiPh-based SMA.To address these challenges, we propose SOFTONIC, a first-of-its-kind photonic SMA architecture designed for ultra-high energy efficiency and speedup in Transformer acceleration. SOFTONIC integrates a microdisk (MD)-based vector decomposition unit, a novel photonic polynomial unit with a range reduction technique, and a unique SiPh lookup table. Simulations of SOFTONIC using industry-standard CAD tools and AI workloads demonstrate a (i) 109× improvement in latency; (ii) an 80% reduction in power consumption; (iii) upto 199 × higher compute-density compared to leading digital and analog Softmax hardware solutions. These improvements come with an area overhead of 0.0157 mm2, highlighting the potential of SOFTONIC for efficient, large-scale photonic AI acceleration.
Author supplied keywords
Cite
CITATION STYLE
Dash, P., Jiang, A., & Dang, D. (2025). SOFTONIC: A Photonic Design Approach to Softmax Activation for High-Speed Fully Analog AI Acceleration. In Proceedings of the ACM Great Lakes Symposium on VLSI, GLSVLSI (pp. 118–125). Association for Computing Machinery. https://doi.org/10.1145/3716368.3735220
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.