EdgeCrafter: Compact ViTs for Edge Dense Prediction via Task-Specialized Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Longfei, Hou, Yongjie, Li, Yang, Wang, Qirui, Sha, Youyang, Yu, Yongjun, Wang, Yinzhi, Ru, Peizhe, Yu, Xuanlong, Shen, Xi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911547872247808
author Liu, Longfei
Hou, Yongjie
Li, Yang
Wang, Qirui
Sha, Youyang
Yu, Yongjun
Wang, Yinzhi
Ru, Peizhe
Yu, Xuanlong
Shen, Xi
author_facet Liu, Longfei
Hou, Yongjie
Li, Yang
Wang, Qirui
Sha, Youyang
Yu, Yongjun
Wang, Yinzhi
Ru, Peizhe
Yu, Xuanlong
Shen, Xi
contents Deploying high-performance dense prediction models on resource-constrained edge devices remains challenging due to strict limits on computation and memory. In practice, lightweight systems for object detection, instance segmentation, and pose estimation are still dominated by CNN-based architectures such as YOLO, while compact Vision Transformers (ViTs) often struggle to achieve similarly strong accuracy efficiency tradeoff, even with large scale pretraining. We argue that this gap is largely due to insufficient task specific representation learning in small scale ViTs, rather than an inherent mismatch between ViTs and edge dense prediction. To address this issue, we introduce EdgeCrafter, a unified compact ViT framework for edge dense prediction centered on ECDet, a detection model built from a distilled compact backbone and an edge-friendly encoder decoder design. On the COCO dataset, ECDet-S achieves 51.7 AP with fewer than 10M parameters using only COCO annotations. For instance segmentation, ECInsSeg achieves performance comparable to RF-DETR while using substantially fewer parameters. For pose estimation, ECPose-X reaches 74.8 AP, significantly outperforming YOLO26Pose-X (71.6 AP). These results show that compact ViTs, when paired with task-specialized distillation and edge-aware design, can be a practical and competitive option for edge dense prediction. Code is available at: https://intellindust-ai-lab.github.io/projects/EdgeCrafter/
format Preprint
id arxiv_https___arxiv_org_abs_2603_18739
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EdgeCrafter: Compact ViTs for Edge Dense Prediction via Task-Specialized Distillation
Liu, Longfei
Hou, Yongjie
Li, Yang
Wang, Qirui
Sha, Youyang
Yu, Yongjun
Wang, Yinzhi
Ru, Peizhe
Yu, Xuanlong
Shen, Xi
Computer Vision and Pattern Recognition
Deploying high-performance dense prediction models on resource-constrained edge devices remains challenging due to strict limits on computation and memory. In practice, lightweight systems for object detection, instance segmentation, and pose estimation are still dominated by CNN-based architectures such as YOLO, while compact Vision Transformers (ViTs) often struggle to achieve similarly strong accuracy efficiency tradeoff, even with large scale pretraining. We argue that this gap is largely due to insufficient task specific representation learning in small scale ViTs, rather than an inherent mismatch between ViTs and edge dense prediction. To address this issue, we introduce EdgeCrafter, a unified compact ViT framework for edge dense prediction centered on ECDet, a detection model built from a distilled compact backbone and an edge-friendly encoder decoder design. On the COCO dataset, ECDet-S achieves 51.7 AP with fewer than 10M parameters using only COCO annotations. For instance segmentation, ECInsSeg achieves performance comparable to RF-DETR while using substantially fewer parameters. For pose estimation, ECPose-X reaches 74.8 AP, significantly outperforming YOLO26Pose-X (71.6 AP). These results show that compact ViTs, when paired with task-specialized distillation and edge-aware design, can be a practical and competitive option for edge dense prediction. Code is available at: https://intellindust-ai-lab.github.io/projects/EdgeCrafter/
title EdgeCrafter: Compact ViTs for Edge Dense Prediction via Task-Specialized Distillation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.18739