Abstract:A cost-effective multi-tenant neural network execution is becoming one of the most important design goals for modern neural network accelerators. For example, as emerging AI services consist of many heterogeneous neural network executions, a cloud provider wants to serve a large number of clients using a single AI accelerator for improving its cost effectiveness. Therefore, an ideal next-generation neural network accelerator should support a simultaneous multi-neural network execution, while fully utilizing its hardware resources. However, existing accelerators which are optimized for a single neural network execution can suffer from severe resource underutilization when running multiple neural networks, mainly due to the load imbalance between computation and memory-access tasks from different neural networks. In this paper, we propose AI-MultiTasking (AI-MT), a novel accelerator architecture which enables a cost-effective, high-performance multi-neural network execution. The key idea of AI-MT is to fully utilize the accelerator's computation resources and memory bandwidth by matching compute- and memory-intensive tasks from different networks and executing them in parallel. However, it is highly challenging to find and schedule the best load-matching tasks from different neural networks during runtime, without significantly increasing the size of on-chip memory. To overcome the challenges, AI-MT first creates fine-grain tasks at compile time by dividing each layer into multiple identical sub-layers. During runtime, AI-MT dynamically applies three sub-layer scheduling methods: memory block prefetching and compute block merging for the best resource load matching, and memory block eviction for the minimum on-chip memory footprint. Our evaluations using MLPerf benchmarks show that AI-MT achieves up to 1.57x speedup over the baseline scheduling method.

Multi-Objective Hardware-Mapping Co-Optimisation for Multi-DNN Workloads on Chiplet-based Accelerators

Optimizing DNN Inference on Multi-Accelerator SoCs at Training-time

MOC: Multi-Objective Mobile CPU-GPU Co-Optimization for Power-Efficient DNN Inference

NN-Baton: DNN Workload Orchestration and Chiplet Granularity Exploration for Multichip Accelerators

SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module Accelerators

M2M: A Fine-Grained Mapping Framework to Accelerate Multiple DNNs on a Multi-Chiplet Architecture

Automatic Mapping of Heterogeneous DNN Models on Adaptive Multi-Accelerator Systems

MoCA: Memory-Centric, Adaptive Execution for Multi-Tenant Deep Neural Networks

A Multi-Neural Network Acceleration Architecture

MARS: Exploiting Multi-Level Parallelism for DNN Workloads on Adaptive Multi-Accelerator Systems

Memory and Computation Coordinated Mapping of DNNs Onto Complex Heterogeneous SoC.

Inter-Layer Scheduling Space Exploration for Multi-model Inference on Heterogeneous Chiplets

Towards Efficient IMC Accelerator Design Through Joint Hardware-Workload Co-optimization

Efficient Mapping Space Exploration for a Reconfigurable Neural Accelerator

Model-Platform Optimized Deep Neural Network Accelerator Generation Through Mixed-Integer Geometric Programming.

Multi-Objective Hardware Aware Neural Architecture Search using Hardware Cost Diversity

High-performance application mapping in network-on-chip-based multicore systems

INDM: Chiplet-Based Interconnect Network and Dataflow Mapping for DNN Accelerators

Software-defined Design Space Exploration for an Efficient DNN Accelerator Architecture

Mrna: Enabling Efficient Mapping Space Exploration for a Reconfiguration Neural Accelerator

Monad: Towards Cost-effective Specialization for Chiplet-based Spatial Accelerators