Hardware–Software Co-Design for Real-Time Latency–Accuracy Navigation in Tiny Machine Learning Applications

Behnam, Payman; Tong, Jianming; Khare, Alind; Chen, Yangyu; Pan, Yue; Gadikar, Pranav; Bambhaniya, Abhimanyu; Krishna, Tushar; Tumanov, Alexey

doi:10.1109/MM.2023.3317243

Citation Details

Hardware–Software Co-Design for Real-Time Latency–Accuracy Navigation in Tiny Machine Learning Applications

Tiny machine learning (TinyML) applications increasingly operate in dynamically changing deployment scenarios, requiring optimization for both accuracy and latency. Existing methods mainly target a single point in the accuracy/latency tradeoff space, which is insufficient as no single static point can be optimal under variable conditions. We draw on a recently proposed weight-shared SuperNet mechanism to enable serving a stream of queries that activates different SubNets within a SuperNet. This creates an opportunity to exploit the inherent temporal locality of different queries that use the same SuperNet. We propose a hardware–software co-design called SUSHI that introduces a novel SubGraph Stationary optimization. SUSHI consists of a novel field-programmable gate array implementation and a software scheduler that controls which SubNets to serve and which SubGraph to cache in real time. SUSHI yields up to a 32% improvement in latency, 0.98% increase in served accuracy, and achieves up to 78.7% off-chip energy saved across several neural network architectures. more »

Award ID(s):: 2029004

PAR ID:: 10510165

Author(s) / Creator(s):: Behnam, Payman; Tong, Jianming; Khare, Alind; Chen, Yangyu; Pan, Yue; Gadikar, Pranav; Bambhaniya, Abhimanyu; Krishna, Tushar; Tumanov, Alexey

Publisher / Repository:: IEEE Computer Society

Date Published:: 2023-11-01

Journal Name:: IEEE Micro

Volume:: 43

Issue:: 6

ISSN:: 0272-1732

Page Range / eLocation ID:: 93 to 101

Subject(s) / Keyword(s):: Training, Real Time Systems, Optimization, Neural Networks, System On Chip, Software, Tiny Machine Learning, Machine Learning, Microcontrollers, Software Design, Machine Learning Applications, Tiny Machine Learning, Neural Network, Deep Neural Network, Search Space, Autonomous Vehicles, Design Space, Self Driving, Temporal Localization, Caching, Data Reuse, Neural Architecture Search, Deployment Phase, Improvement In Latency, Blue Dots, Dot Product, Load Data, Latency Reduction, Cache Hit, Shared Weights

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Journal Article:
https://doi.org/10.1109/MM.2023.3317243

More Like this