MicroViTv2: Beyond the FLOPS for Edge Energy-Friendly Vision Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Setyawan, Novendra, Sun, Chi-Chia, Hsu, Mao-Hsiu, Kuo, Wen-Kai, Hsieh, Jun-Wei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MicroViT: A Vision Transformer with Low Complexity Self Attention for Edge Device
by: Setyawan, Novendra, et al.
Published: (2025)
by: Setyawan, Novendra, et al.
Published: (2025)
FaceLiVT: Face Recognition using Linear Vision Transformer with Structural Reparameterization For Mobile Device
by: Setyawan, Novendra, et al.
Published: (2025)
by: Setyawan, Novendra, et al.
Published: (2025)
FaceLiVTv2: An Improved Hybrid Architecture for Efficient Mobile Face Recognition
by: Setyawan, Novendra, et al.
Published: (2026)
by: Setyawan, Novendra, et al.
Published: (2026)
Fast-COS: A Fast One-Stage Object Detector Based on Reparameterized Attention Vision Transformer for Autonomous Driving
by: Setyawan, Novendra, et al.
Published: (2025)
by: Setyawan, Novendra, et al.
Published: (2025)
ParFormer: A Vision Transformer with Parallel Mixer and Sparse Channel Attention Patch Embedding
by: Setyawan, Novendra, et al.
Published: (2024)
by: Setyawan, Novendra, et al.
Published: (2024)
Fast Person Detection Using YOLOX With AI Accelerator For Train Station Safety
by: Achmadiah, Mas Nurul, et al.
Published: (2026)
by: Achmadiah, Mas Nurul, et al.
Published: (2026)
Energy-Efficient Fast Object Detection on Edge Devices for IoT Systems
by: Achmadiah, Mas Nurul, et al.
Published: (2026)
by: Achmadiah, Mas Nurul, et al.
Published: (2026)
RepSFNet : A Single Fusion Network with Structural Reparameterization for Crowd Counting
by: Achmadiah, Mas Nurul, et al.
Published: (2026)
by: Achmadiah, Mas Nurul, et al.
Published: (2026)
A Beginner-Friendly ESP32-Based Remote Control Application
by: Setyawan, Herlin
Published: (2026)
by: Setyawan, Herlin
Published: (2026)
FiTv2: Scalable and Improved Flexible Vision Transformer for Diffusion Model
by: Wang, ZiDong, et al.
Published: (2024)
by: Wang, ZiDong, et al.
Published: (2024)
ADFQ-ViT: Activation-Distribution-Friendly Post-Training Quantization for Vision Transformers
by: Jiang, Yanfeng, et al.
Published: (2024)
by: Jiang, Yanfeng, et al.
Published: (2024)
Symmetry-Enforced Quadratic Degradability Beyond Low Dimensions
by: Lo, Yun-Feng, et al.
Published: (2024)
by: Lo, Yun-Feng, et al.
Published: (2024)
StrucTexTv3: An Efficient Vision-Language Model for Text-rich Image Perception, Comprehension, and Beyond
by: Lyu, Pengyuan, et al.
Published: (2024)
by: Lyu, Pengyuan, et al.
Published: (2024)
FLOPS: Forward Learning with OPtimal Sampling
by: Ren, Tao, et al.
Published: (2024)
by: Ren, Tao, et al.
Published: (2024)
DopQ-ViT: Towards Distribution-Friendly and Outlier-Aware Post-Training Quantization for Vision Transformers
by: Yang, Lianwei, et al.
Published: (2024)
by: Yang, Lianwei, et al.
Published: (2024)
Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection
by: Huang, Xian-Hong, et al.
Published: (2025)
by: Huang, Xian-Hong, et al.
Published: (2025)
Oracle Separation between Noisy Quantum Polynomial Time and the Polynomial Hierarchy
by: Chia, Nai-Hui, et al.
Published: (2024)
by: Chia, Nai-Hui, et al.
Published: (2024)
Accelerating Maximal Biclique Enumeration on GPUs
by: Hsieh, Chou-Ying, et al.
Published: (2024)
by: Hsieh, Chou-Ying, et al.
Published: (2024)
On the Capacity of Zero-Drift First Arrival Position Channels in Diffusive Molecular Communication
by: Lee, Yen-Chi, et al.
Published: (2022)
by: Lee, Yen-Chi, et al.
Published: (2022)
ButterflyViT: 354$\times$ Expert Compression for Edge Vision Transformers
by: Karmore, Aryan
Published: (2026)
by: Karmore, Aryan
Published: (2026)
VAT: Vision Action Transformer by Unlocking Full Representation of ViT
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
Quantum Walks on Simplicial Complexes and Harmonic Homology: Application to Topological Data Analysis with Superpolynomial Speedups
by: Hayakawa, Ryu, et al.
Published: (2024)
by: Hayakawa, Ryu, et al.
Published: (2024)
Qsyn: A Developer-Friendly Quantum Circuit Synthesis Framework for NISQ Era and Beyond
by: Lau, Mu-Te, et al.
Published: (2024)
by: Lau, Mu-Te, et al.
Published: (2024)
Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the Edge
by: Dumoulin, Joren, et al.
Published: (2025)
by: Dumoulin, Joren, et al.
Published: (2025)
LL-ViT: Edge Deployable Vision Transformers with Look Up Table Neurons
by: Nag, Shashank, et al.
Published: (2025)
by: Nag, Shashank, et al.
Published: (2025)
An Alternative to FLOPS Regularization to Effectively Productionize SPLADE-Doc
by: Porco, Aldo, et al.
Published: (2025)
by: Porco, Aldo, et al.
Published: (2025)
LocalViT: Analyzing Locality in Vision Transformers
by: Li, Yawei, et al.
Published: (2021)
by: Li, Yawei, et al.
Published: (2021)
I've Got 99 Problems But FLOPS Ain't One
by: Gherghescu, Alexandru M., et al.
Published: (2024)
by: Gherghescu, Alexandru M., et al.
Published: (2024)
Enhancing Learnable Descriptive Convolutional Vision Transformer for Face Anti-Spoofing
by: Huanga, Pei-Kai, et al.
Published: (2025)
by: Huanga, Pei-Kai, et al.
Published: (2025)
Block-based Symmetric Pruning and Fusion for Efficient Vision Transformers
by: Hsieh, Yi-Kuan, et al.
Published: (2025)
by: Hsieh, Yi-Kuan, et al.
Published: (2025)
EA-ViT: Efficient Adaptation for Elastic Vision Transformer
by: Zhu, Chen, et al.
Published: (2025)
by: Zhu, Chen, et al.
Published: (2025)
BDS ansatz in ABJM via scaffolding triangulations
by: Huang, Yu-tin, et al.
Published: (2025)
by: Huang, Yu-tin, et al.
Published: (2025)
ModeTv2: GPU-accelerated Motion Decomposition Transformer for Pairwise Optimization in Medical Image Registration
by: Wang, Haiqiao, et al.
Published: (2024)
by: Wang, Haiqiao, et al.
Published: (2024)
CanvOI, an Oncology Intelligence Foundation Model: Scaling FLOPS Differently
by: Zalach, Jonathan, et al.
Published: (2024)
by: Zalach, Jonathan, et al.
Published: (2024)
Digital Workflow and Guided Surgery in Implant Therapy—Literature Review and Practical Tips to Optimize Precision
by: Chia‐Sheng Chen, et al.
Published: (2025)
by: Chia‐Sheng Chen, et al.
Published: (2025)
Trio-ViT: Post-Training Quantization and Acceleration for Softmax-Free Efficient Vision Transformer
by: Shi, Huihong, et al.
Published: (2024)
by: Shi, Huihong, et al.
Published: (2024)
Gudder–Nagy’s Theorem for Hilbert K ( H )‐Modules
by: Ming-Hsiu Hsu
Published: (2025)
by: Ming-Hsiu Hsu
Published: (2025)
Seg the HAB: Language-Guided Geospatial Algae Bloom Reasoning and Segmentation
by: Hsieh, Patterson, et al.
Published: (2025)
by: Hsieh, Patterson, et al.
Published: (2025)
Invited Paper: BitMedViT: Ternary-Quantized Vision Transformer for Medical AI Assistants on the Edge
by: Walczak, Mikolaj, et al.
Published: (2025)
by: Walczak, Mikolaj, et al.
Published: (2025)
Transform Arbitrary Good Quantum LDPC Codes into Good Geometrically Local Codes in Any Dimension
by: Li, Xingjian, et al.
Published: (2024)
by: Li, Xingjian, et al.
Published: (2024)
Similar Items
-
MicroViT: A Vision Transformer with Low Complexity Self Attention for Edge Device
by: Setyawan, Novendra, et al.
Published: (2025) -
FaceLiVT: Face Recognition using Linear Vision Transformer with Structural Reparameterization For Mobile Device
by: Setyawan, Novendra, et al.
Published: (2025) -
FaceLiVTv2: An Improved Hybrid Architecture for Efficient Mobile Face Recognition
by: Setyawan, Novendra, et al.
Published: (2026) -
Fast-COS: A Fast One-Stage Object Detector Based on Reparameterized Attention Vision Transformer for Autonomous Driving
by: Setyawan, Novendra, et al.
Published: (2025) -
ParFormer: A Vision Transformer with Parallel Mixer and Sparse Channel Attention Patch Embedding
by: Setyawan, Novendra, et al.
Published: (2024)