Abstract:Homomorphic encryption is one of the main solutions for building secure and privacy-preserving solutions for Machine Learning as a Service. This motivates the development of homomorphic algorithms for the main building blocks of AI, typically for the components of the various types of neural networks architectures. Among those components, we focus on the Softmax function, defined by $\mathrm{SM}(\mathbf{x}) = \left(\exp(x_i) / \sum_{j=1}^n \exp(x_j) \right)_{1\le i\le n}$. This function is deemed to be one of the most difficult to evaluate homomorphically, because of its multivariate nature and of the very large range of values for $\exp(x_i)$. The available homomorphic algorithms remain restricted, especially in large dimensions, while important applications such as Large Language Models (LLM) require computing Softmax over large dimensional vectors. In terms of multiplicative depth of the computation (a suitable measure of cost for homomorphic algorithms), our algorithm achieves $O(\log n)$ complexity for a fixed range of inputs, where $n$ is the Softmax dimension. Our algorithm is especially adapted to the situation where we must compute many Softmax at the same time, for instance, in the LLM situation. In that case, assuming that all Softmax calls are packed into $m$ ciphtertexts, the asymptotic amortized multiplicative depth cost per ciphertext is, again over a fixed range, $O(1 + m/N)$ for $N$ the homomorphic ring degree. The main ingredient of our algorithms is a normalize-and-square strategy, which interlaces the exponential computation over a large range and normalization, decomposing both in stabler and cheaper smaller steps. Comparing ourselves to the state of the art, our experiments show, in practice, a good accuracy and a gain of a factor 2.5 to 8 compared to state of the art solutions.

Hardware-Efficient SoftMax Architecture With Bit-Wise Exponentiation and Reciprocal Calculation

Efficient Hardware Architecture of Softmax Layer in Deep Neural Network

High-Precision Method and Architecture for Base-2 Softmax Function in DNN Training.

Hardware-Aware Softmax Approximation for Deep Neural Networks

A High Speed SoftMax VLSI Architecture Based on Basic-Split

Base-2 Softmax Function: Suitability for Training and Efficient Hardware Implementation

ConSmax: Hardware-Friendly Alternative Softmax with Learnable Parameters

Hyft: A Reconfigurable Softmax Accelerator with Hybrid Numeric Format for both Training and Inference

SoftmAP: Software-Hardware Co-design for Integer-Only Softmax on Associative Processors

Efficient FPGA Implementation of softmax Layer in Deep Neural Network

SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference

Design Space Exploration of Neural Network Activation Function Circuits

2 β-softmax: A Hardware-Friendly Activation Function with Low Complexity and High Performance

Softmax Optimizations for Intel Xeon Processor-based Platforms

Optimization of Softmax Layer in Deep Neural Network Using Integral Stochastic Computation

Towards Ultra-High Performance and Energy Efficiency of Deep Learning Systems: An Algorithm-Hardware Co-Optimization Framework

Hardware Platform-Aware Binarized Neural Network Model Optimization

Fast and Accurate Homomorphic Softmax Evaluation

A High Performance Reconfigurable Hardware Architecture for Lightweight Convolutional Neural Network

Neural Synaptic Plasticity-Inspired Computing: A High Computing Efficient Deep Convolutional Neural Network Accelerator

A Flexible Template for Edge Generative AI with High-Accuracy Accelerated Softmax & GELU