NuWa: Deriving Lightweight Task-Specific Vision Transformers for Edge Devices
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Ziteng, He, Qiang, Li, Bing, Chen, Feifei, Yang, Yun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Partitioning Vision Transformer on Edge Devices for Distributed Inference
by: Liu, Xiang, et al.
Published: (2024)
by: Liu, Xiang, et al.
Published: (2024)
Lightweight Optimal-Transport Harmonization on Edge Devices
by: Larchenko, Maria, et al.
Published: (2025)
by: Larchenko, Maria, et al.
Published: (2025)
Lightweight Remote Sensing Scene Classification on Edge Devices via Knowledge Distillation and Early-exit
by: Zhao, Yang, et al.
Published: (2025)
by: Zhao, Yang, et al.
Published: (2025)
How Lightweight Can A Vision Transformer Be
by: Tan, Jen Hong
Published: (2024)
by: Tan, Jen Hong
Published: (2024)
MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
LVLM-Aided Alignment of Task-Specific Vision Models
by: Koebler, Alexander, et al.
Published: (2025)
by: Koebler, Alexander, et al.
Published: (2025)
A Lightweight Dual-Branch System for Weakly-Supervised Video Anomaly Detection on Consumer Edge Devices
by: Jiang, Wen-Dong, et al.
Published: (2024)
by: Jiang, Wen-Dong, et al.
Published: (2024)
LWMSCNN-SE: A Lightweight Multi-Scale Network for Efficient Maize Disease Classification on Edge Devices
by: Weloday, Fikadu, et al.
Published: (2026)
by: Weloday, Fikadu, et al.
Published: (2026)
Task-Specific Adaptation of Segmentation Foundation Model via Prompt Learning
by: Kim, Hyung-Il, et al.
Published: (2024)
by: Kim, Hyung-Il, et al.
Published: (2024)
TAP-ViTs: Task-Adaptive Pruning for On-Device Deployment of Vision Transformers
by: Wang, Zhibo, et al.
Published: (2026)
by: Wang, Zhibo, et al.
Published: (2026)
Token Pruning using a Lightweight Background Aware Vision Transformer
by: Sah, Sudhakar, et al.
Published: (2024)
by: Sah, Sudhakar, et al.
Published: (2024)
Dynamic Integration of Task-Specific Adapters for Class Incremental Learning
by: Li, Jiashuo, et al.
Published: (2024)
by: Li, Jiashuo, et al.
Published: (2024)
EdgeDiT: Hardware-Aware Diffusion Transformers for Efficient On-Device Image Generation
by: Kodavanti, Sravanth, et al.
Published: (2026)
by: Kodavanti, Sravanth, et al.
Published: (2026)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
by: Tian, Kexin, et al.
Published: (2025)
by: Tian, Kexin, et al.
Published: (2025)
ButterflyViT: 354$\times$ Expert Compression for Edge Vision Transformers
by: Karmore, Aryan
Published: (2026)
by: Karmore, Aryan
Published: (2026)
Can Vision-Language Models Replace Human Annotators: A Case Study with CelebA Dataset
by: Lu, Haoming, et al.
Published: (2024)
by: Lu, Haoming, et al.
Published: (2024)
SAC-ViT: Semantic-Aware Clustering Vision Transformer with Early Exit
by: Hu, Youbing, et al.
Published: (2025)
by: Hu, Youbing, et al.
Published: (2025)
Characterizing Disparity Between Edge Models and High-Accuracy Base Models for Vision Tasks
by: Wang, Zhenyu, et al.
Published: (2024)
by: Wang, Zhenyu, et al.
Published: (2024)
EdgeEar: Efficient and Accurate Ear Recognition for Edge Devices
by: Lendering, Camile, et al.
Published: (2025)
by: Lendering, Camile, et al.
Published: (2025)
Edge-Enhanced Vision Transformer Framework for Accurate AI-Generated Image Detection
by: Das, Dabbrata, et al.
Published: (2025)
by: Das, Dabbrata, et al.
Published: (2025)
Towards Optimal Trade-offs in Knowledge Distillation for CNNs and Vision Transformers at the Edge
by: Violos, John, et al.
Published: (2024)
by: Violos, John, et al.
Published: (2024)
ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding
by: Hsieh, ZongHan, et al.
Published: (2025)
by: Hsieh, ZongHan, et al.
Published: (2025)
Enhancing Vision-Language Models for Autonomous Driving through Task-Specific Prompting and Spatial Reasoning
by: Wu, Aodi, et al.
Published: (2025)
by: Wu, Aodi, et al.
Published: (2025)
MobileUNETR: A Lightweight End-To-End Hybrid Vision Transformer For Efficient Medical Image Segmentation
by: Perera, Shehan, et al.
Published: (2024)
by: Perera, Shehan, et al.
Published: (2024)
Efficient Adaptation of Pre-trained Vision Transformer via Householder Transformation
by: Dong, Wei, et al.
Published: (2024)
by: Dong, Wei, et al.
Published: (2024)
Early Explorations of Lightweight Models for Wound Segmentation on Mobile Devices
by: Borst, Vanessa, et al.
Published: (2024)
by: Borst, Vanessa, et al.
Published: (2024)
Reciprocal Attention Mixing Transformer for Lightweight Image Restoration
by: Choi, Haram, et al.
Published: (2023)
by: Choi, Haram, et al.
Published: (2023)
Edge-AI for Agriculture: Lightweight Vision Models for Disease Detection in Resource-Limited Settings
by: Joshi, Harsh
Published: (2024)
by: Joshi, Harsh
Published: (2024)
EfficientFSL: Enhancing Few-Shot Classification via Query-Only Tuning in Vision Transformers
by: Liao, Wenwen, et al.
Published: (2026)
by: Liao, Wenwen, et al.
Published: (2026)
Deep Extrinsic Manifold Representation for Vision Tasks
by: Zhang, Tongtong, et al.
Published: (2024)
by: Zhang, Tongtong, et al.
Published: (2024)
Proximal Vision Transformer: Enhancing Feature Representation through Two-Stage Manifold Geometry
by: Yun, Haoyu, et al.
Published: (2025)
by: Yun, Haoyu, et al.
Published: (2025)
EdgeNAT: Transformer for Efficient Edge Detection
by: Jie, Jinghuai, et al.
Published: (2024)
by: Jie, Jinghuai, et al.
Published: (2024)
LF-ViT: Reducing Spatial Redundancy in Vision Transformer for Efficient Image Recognition
by: Hu, Youbing, et al.
Published: (2024)
by: Hu, Youbing, et al.
Published: (2024)
On-Demand Multi-Task Sparsity for Efficient Large-Model Deployment on Edge Devices
by: Huang, Lianming, et al.
Published: (2025)
by: Huang, Lianming, et al.
Published: (2025)
EDT: An Efficient Diffusion Transformer Framework Inspired by Human-like Sketching
by: Chen, Xinwang, et al.
Published: (2024)
by: Chen, Xinwang, et al.
Published: (2024)
Lightweight Distillation of SAM 3 and DINOv3 for Edge-Deployable Individual-Level Livestock Monitoring and Longitudinal Visual Analytics
by: Yang, Haiyu, et al.
Published: (2026)
by: Yang, Haiyu, et al.
Published: (2026)
Sensitive Image Classification by Vision Transformers
by: He, Hanxian, et al.
Published: (2024)
by: He, Hanxian, et al.
Published: (2024)
A 2D Semantic-Aware Position Encoding for Vision Transformers
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
Real-Time Pedestrian Detection on IoT Edge Devices: A Lightweight Deep Learning Approach
by: Alfikri, Muhammad Dany, et al.
Published: (2024)
by: Alfikri, Muhammad Dany, et al.
Published: (2024)
Lightweight Multimodal Adaptation of Vision Language Models for Species Recognition and Habitat Context Interpretation in Drone Thermal Imagery
by: Chen, Hao, et al.
Published: (2026)
by: Chen, Hao, et al.
Published: (2026)
Similar Items
-
Efficient Partitioning Vision Transformer on Edge Devices for Distributed Inference
by: Liu, Xiang, et al.
Published: (2024) -
Lightweight Optimal-Transport Harmonization on Edge Devices
by: Larchenko, Maria, et al.
Published: (2025) -
Lightweight Remote Sensing Scene Classification on Edge Devices via Knowledge Distillation and Early-exit
by: Zhao, Yang, et al.
Published: (2025) -
How Lightweight Can A Vision Transformer Be
by: Tan, Jen Hong
Published: (2024) -
MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning
by: Liu, Yi, et al.
Published: (2025)