A Transverse-Read-assisted Fast Valid-Bits Collection in Stochastic Computing MACs for Energy-Efficient in-RTM DNNs
Jihe Wang,Zhiying Zhang,Xingwu Dong,Danghui Wang
2024-10-22
Abstract:It looks very attractive to coordinate racetrack-memory (RM) and stochastic-computing (SC) jointly to build an ultra-low power <a class="link-external link-http" href="http://neuron-architecture.However" rel="external noopener nofollow">this http URL</a>, the above combination has always been questioned in a fatal weakness that the heavy valid-bits collection of RM-MTJ, a.k.a. accumulative parallel counters (APCs), cannot physically match the requirement for energy-efficient in-memory <a class="link-external link-http" href="http://DNNs.Fortunately" rel="external noopener nofollow">this http URL</a>, a recently developed Transverse-Read (TR) provides a lightweight collection of valid-bits by detecting domain-wall resistance between a couple of MTJs on a single <a class="link-external link-http" href="http://nanowire.In" rel="external noopener nofollow">this http URL</a> this work, we first propose a neuron-architecture that utilizes parallel TRs to build an ultra-fast valid-bits collection for SC, in which, a vector multiplication is successfully degraded as swift <a class="link-external link-http" href="http://TRs.To" rel="external noopener nofollow">this http URL</a> solve huge storage for full stochastic sequences caused by the limited TR banks, a hybrid coding, pseudo-fractal compression, is designed to generate stochastic sequences by <a class="link-external link-http" href="http://segments.To" rel="external noopener nofollow">this http URL</a> overcome the misalignment by the parallel early-termination, an asynchronous schedule of TR is further designed to regularize the vectorization, in which, the valid-bits from different lanes are merged in multiple RM-stacks for vector-level valid-bits <a class="link-external link-http" href="http://collection.However" rel="external noopener nofollow">this http URL</a>, an inherent defect of TR, i.e., neighbor parts cannot be accessed simultaneously, could limit the throughput of the parallel vector multiplication, therefore, an interleaving data placement is used for full utilization of memory bus among different <a class="link-external link-http" href="http://vectors.The" rel="external noopener nofollow">this http URL</a> results show that the SC-MAC assisted with TR achieves $2.88\times-4.40\times $speedup compared to CORUSCANT, at the same time, energy consumption is reduced by $1.26\times-1.42\times$.
Distributed, Parallel, and Cluster Computing