Abstract:Machine intelligence, especially using convolutional neural networks (CNNs), has become a large area of research over the past years. Increasingly sophisticated hardware accelerators are proposed that exploit e.g. the sparsity in computations and make use of reduced precision arithmetic to scale down the energy consumption. However, future platforms require more than just energy efficiency: Scalability is becoming an increasingly important factor. The required effort for physical implementation grows with the size of the accelerator making it more difficult to meet target constraints. Using many-core platforms consisting of several homogeneous cores can alleviate the aforementioned limitations with regard to physical implementation at the expense of an increased dataflow mapping effort. While the dataflow in CNNs is deterministic and can therefore be optimized offline, the problem of finding a suitable scheme that minimizes both runtime and off-chip memory accesses is a challenging task which becomes even more complex if an interconnect system is involved. This work presents an automated mapping strategy starting at the single-core level with different optimization targets for minimal runtime and minimal off-chip memory accesses. The strategy is then extended towards a suitable many-core mapping scheme and evaluated using a scalable system-level simulation with a network-on-chip interconnect. Design space exploration is performed by mapping the well-known CNNs AlexNet and VGG-16 to platforms of different core counts and computational power per core in order to investigate the trade-offs. Our mapping strategy and system setup is scaled starting from the single core level up to 128 cores, thereby showing the limits of the selected approach.

AOME: Autonomous Optimal Mapping Exploration Using Reinforcement Learning for NoC-based Accelerators Running Neural Networks

An Efficient Algorithm for Mapping Deep Learning Applications on the NoC Architecture

Automatic Generation and Optimization Framework of NoC-Based Neural Network Accelerator Through Reinforcement Learning

Optimized Mapping Spiking Neural Networks onto Network-on-Chip.

Mapping Very Large Scale Spiking Neuron Network to Neuromorphic Hardware.

Approaching the mapping limit with closed-loop mapping strategy for deploying neural networks on neuromorphic hardware

High-performance application mapping in network-on-chip-based multicore systems

A Multi-Objective Model Oriented Mapping Approach for NoC-based Computing Systems

A Load Balanced Mapping for Spiking Neural Network

An optimized mapping algorithm based on simulated annealing for regular NoC architecture

M2M: A Fine-Grained Mapping Framework to Accelerate Multiple DNNs on a Multi-Chiplet Architecture

Incremental Run-time Application Mapping for Heterogeneous Network on Chip

A Mapping Model of SNNs to Neuromorphic Hardware

Fast-OverlaPIM: A Fast Overlap-driven Mapping Framework for Processing In-Memory Neural Network Acceleration

ARES: A Mapping Framework of DNNs Towards Diverse PIMs with General Abstractions

Memory and Computation Coordinated Mapping of DNNs Onto Complex Heterogeneous SoC.

NN-Baton: DNN Workload Orchestration and Chiplet Granularity Exploration for Multichip Accelerators

An Efficient Task Mapping Algorithm with Power-Aware Optimization for Network on Chip

TTNNM: Thermal- and Traffic-Aware Neural Network Mapping on 3D-Noc-based Accelerator

Latency-aware mapping for 3D NoC using rank-based multi-objective genetic algorithm

Dataflow Aware Mapping of Convolutional Neural Networks Onto Many-Core Platforms With Network-on-Chip Interconnect