Image Recognition with Online Lightweight Vision Transformer: A Survey
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zherui, Xu, Rongtao, Zhou, Jie, Wang, Changwei, Pei, Xingtian, Xu, Wenhao, Zhang, Jiguang, Guo, Li, Gao, Longxiang, Xu, Wenbo, Xu, Shibiao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FDBPL: Faster Distillation-Based Prompt Learning for Region-Aware Vision-Language Models Adaptation
by: Zhang, Zherui, et al.
Published: (2025)
by: Zhang, Zherui, et al.
Published: (2025)
DeliCIR: Deliberative Test-Time Evolutionary Hierarchical Multi-Agents for Composed Image Retrieval
by: Pei, Xingtian, et al.
Published: (2026)
by: Pei, Xingtian, et al.
Published: (2026)
SAMamba: Adaptive State Space Modeling with Hierarchical Vision for Infrared Small Target Detection
by: Xu, Wenhao, et al.
Published: (2025)
by: Xu, Wenhao, et al.
Published: (2025)
SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place Recognition
by: Chen, Shunpeng, et al.
Published: (2025)
by: Chen, Shunpeng, et al.
Published: (2025)
CAE-DFKD: Bridging the Transferability Gap in Data-Free Knowledge Distillation
by: Zhang, Zherui, et al.
Published: (2025)
by: Zhang, Zherui, et al.
Published: (2025)
SkinFormer: Learning Statistical Texture Representation with Transformer for Skin Lesion Segmentation
by: Xu, Rongtao, et al.
Published: (2024)
by: Xu, Rongtao, et al.
Published: (2024)
Generalization Boosted Adapter for Open-Vocabulary Segmentation
by: Xu, Wenhao, et al.
Published: (2024)
by: Xu, Wenhao, et al.
Published: (2024)
Focus on Local: Finding Reliable Discriminative Regions for Visual Place Recognition
by: Wang, Changwei, et al.
Published: (2025)
by: Wang, Changwei, et al.
Published: (2025)
PSTNet: Enhanced Polyp Segmentation with Multi-scale Alignment and Frequency Domain Integration
by: Xu, Wenhao, et al.
Published: (2024)
by: Xu, Wenhao, et al.
Published: (2024)
Region Matters: Efficient and Reliable Region-Aware Visual Place Recognition
by: Chen, Shunpeng, et al.
Published: (2026)
by: Chen, Shunpeng, et al.
Published: (2026)
Spectral Prompt Tuning:Unveiling Unseen Classes for Zero-Shot Semantic Segmentation
by: Xu, Wenhao, et al.
Published: (2023)
by: Xu, Wenhao, et al.
Published: (2023)
HCF-Net: Hierarchical Context Fusion Network for Infrared Small Object Detection
by: Xu, Shibiao, et al.
Published: (2024)
by: Xu, Shibiao, et al.
Published: (2024)
Local Feature Matching Using Deep Learning: A Survey
by: Xu, Shibiao, et al.
Published: (2024)
by: Xu, Shibiao, et al.
Published: (2024)
CurriFlow: Curriculum-Guided Depth Fusion with Optical Flow-Based Temporal Alignment for 3D Semantic Scene Completion
by: Lin, Jinzhou, et al.
Published: (2025)
by: Lin, Jinzhou, et al.
Published: (2025)
Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
by: Han, Xiaofeng, et al.
Published: (2025)
by: Han, Xiaofeng, et al.
Published: (2025)
3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering
by: Xu, Rongtao, et al.
Published: (2025)
by: Xu, Rongtao, et al.
Published: (2025)
Segment Anything Model is a Good Teacher for Local Feature Learning
by: Wu, Jingqian, et al.
Published: (2023)
by: Wu, Jingqian, et al.
Published: (2023)
Complementary Information Guided Occupancy Prediction via Multi-Level Representation Fusion
by: Xu, Rongtao, et al.
Published: (2025)
by: Xu, Rongtao, et al.
Published: (2025)
Vision Also You Need: Navigating Out-of-Distribution Detection with Multimodal Large Language Model
by: Xu, Haoran, et al.
Published: (2026)
by: Xu, Haoran, et al.
Published: (2026)
QAGait: Revisit Gait Recognition from a Quality Perspective
by: Wang, Zengbin, et al.
Published: (2024)
by: Wang, Zengbin, et al.
Published: (2024)
Advances in Embodied Navigation Using Large Language Models: A Survey
by: Lin, Jinzhou, et al.
Published: (2023)
by: Lin, Jinzhou, et al.
Published: (2023)
LaplacianFormer:Rethinking Linear Attention with Laplacian Kernel
by: Feng, Zhe, et al.
Published: (2026)
by: Feng, Zhe, et al.
Published: (2026)
Calibration Attention: Learning Reliability-Aware Representations for Vision Transformers
by: Liang, Wenhao, et al.
Published: (2025)
by: Liang, Wenhao, et al.
Published: (2025)
Towards SAR Automatic Target Recognition MultiCategory SAR Image Classification Based on Light Weight Vision Transformer
by: Zhao, Guibin, et al.
Published: (2024)
by: Zhao, Guibin, et al.
Published: (2024)
Human-inspired Global-to-Parallel Multi-scale Encoding for Lightweight Vision Models
by: Xu, Wei
Published: (2026)
by: Xu, Wei
Published: (2026)
S2A: A Unified Framework for Parameter and Memory Efficient Transfer Learning
by: Jin, Tian, et al.
Published: (2025)
by: Jin, Tian, et al.
Published: (2025)
Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction
by: Jia, Mingda, et al.
Published: (2025)
by: Jia, Mingda, et al.
Published: (2025)
UniRestorer: Universal Image Restoration via Adaptively Estimating Image Degradation at Proper Granularity
by: Lin, Jingbo, et al.
Published: (2024)
by: Lin, Jingbo, et al.
Published: (2024)
Fast-Slow Test-Time Adaptation for Online Vision-and-Language Navigation
by: Gao, Junyu, et al.
Published: (2023)
by: Gao, Junyu, et al.
Published: (2023)
Lightweight Vision Transformer with Window and Spatial Attention for Food Image Classification
by: Gao, Xinle, et al.
Published: (2025)
by: Gao, Xinle, et al.
Published: (2025)
Vision Transformers on the Edge: A Comprehensive Survey of Model Compression and Acceleration Strategies
by: Saha, Shaibal, et al.
Published: (2025)
by: Saha, Shaibal, et al.
Published: (2025)
A Lightweight Complex-Valued Deformable CNN for High-Quality Computer-Generated Holography
by: Xie, Shuyang, et al.
Published: (2025)
by: Xie, Shuyang, et al.
Published: (2025)
GSB: Group Superposition Binarization for Vision Transformer with Limited Training Samples
by: Gao, Tian, et al.
Published: (2023)
by: Gao, Tian, et al.
Published: (2023)
Efficient Hyperspectral Image Reconstruction Using Lightweight Separate Spectral Transformers
by: Li, Jianan, et al.
Published: (2026)
by: Li, Jianan, et al.
Published: (2026)
BHViT: Binarized Hybrid Vision Transformer
by: Gao, Tian, et al.
Published: (2025)
by: Gao, Tian, et al.
Published: (2025)
NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation
by: Zhang, Jiazhao, et al.
Published: (2024)
by: Zhang, Jiazhao, et al.
Published: (2024)
UNIT: Unifying Image and Text Recognition in One Vision Encoder
by: Zhu, Yi, et al.
Published: (2024)
by: Zhu, Yi, et al.
Published: (2024)
Task-Aware Dynamic Transformer for Efficient Arbitrary-Scale Image Super-Resolution
by: Xu, Tianyi, et al.
Published: (2024)
by: Xu, Tianyi, et al.
Published: (2024)
MARE: Multimodal Alignment and Reinforcement for Explainable Deepfake Detection via Vision-Language Models
by: Xu, Wenbo, et al.
Published: (2026)
by: Xu, Wenbo, et al.
Published: (2026)
Dual Prompting Image Restoration with Diffusion Transformers
by: Kong, Dehong, et al.
Published: (2025)
by: Kong, Dehong, et al.
Published: (2025)
Similar Items
-
FDBPL: Faster Distillation-Based Prompt Learning for Region-Aware Vision-Language Models Adaptation
by: Zhang, Zherui, et al.
Published: (2025) -
DeliCIR: Deliberative Test-Time Evolutionary Hierarchical Multi-Agents for Composed Image Retrieval
by: Pei, Xingtian, et al.
Published: (2026) -
SAMamba: Adaptive State Space Modeling with Hierarchical Vision for Infrared Small Target Detection
by: Xu, Wenhao, et al.
Published: (2025) -
SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place Recognition
by: Chen, Shunpeng, et al.
Published: (2025) -
CAE-DFKD: Bridging the Transferability Gap in Data-Free Knowledge Distillation
by: Zhang, Zherui, et al.
Published: (2025)