Pruning One More Token is Enough: Leveraging Latency-Workload Non-Linearities for Vision Transformers on the Edge

Eliopoulos, Nick John; Jajal, Purvish; Davis, James; Liu, Gaowen; Thiravathukal, George; Lu, YungHsiang

doi:https://doi.org/10.48550/arXiv.2407.05941

Citation Details

This content will become publicly available on February 28, 2026

Pruning One More Token is Enough: Leveraging Latency-Workload Non-Linearities for Vision Transformers on the Edge

This paper investigates how to efficiently deploy vision transformers on edge devices for small workloads. Recent methods reduce the latency of transformer neural networks by removing or merging tokens, with small accuracy degradation. However, these methods are not designed with edge device deployment in mind: they do not leverage information about the latency-workload trends to improve efficiency. We address this shortcoming in our work. First, we identify factors that affect ViT latency-workload relationships. Second, we determine token pruning schedule by leveraging non-linear latency-workload relationships. Third, we demonstrate a training-free, token pruning method utilizing this schedule. We show other methods may increase latency by 2-30%, while we reduce latency by 9-26%. For similar latency (within 5.2% or 7ms) across devices we achieve 78.6%-84.5% ImageNet1K accuracy, while the state-of-the-art, Token Merging, achieves 45.8%-85.4%. more »

Award ID(s):: 2104709

PAR ID:: 10638751

Author(s) / Creator(s):: Eliopoulos, Nick John; Jajal, Purvish; Davis, James; Liu, Gaowen; Thiravathukal, George; Lu, YungHsiang

Publisher / Repository:: The Computer Vision Foundation.

Date Published:: 2025-02-28

Subject(s) / Keyword(s):: computer vision, token pruning

Format(s):: Medium: X

Location:: Tucson Arizona

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
This content will become publicly available on February 28, 2026
Conference Paper:
https://doi.org/https://doi.org/10.48550/arXiv.2407.05941

More Like this