CoCAViT: Compact Vision Transformer with Robust Global Coordination
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xuyang, Miao, Lingjuan, Zhou, Zhiqiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CAViT -- Channel-Aware Vision Transformer for Dynamic Feature Fusion
by: Safdar, Aon, et al.
Published: (2026)
by: Safdar, Aon, et al.
Published: (2026)
Faster and Better: Reinforced Collaborative Distillation and Self-Learning for Infrared-Visible Image Fusion
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Infrared and Visible Image Fusion with Language-Driven Loss in CLIP Embedding Space
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
Semantic Object-level Modeling for Robust Visual Camera Relocalization
by: Zhu, Yifan, et al.
Published: (2024)
by: Zhu, Yifan, et al.
Published: (2024)
CAViAR: Critic-Augmented Video Agentic Reasoning
by: Menon, Sachit, et al.
Published: (2025)
by: Menon, Sachit, et al.
Published: (2025)
Empirical Recipes for Efficient and Compact Vision-Language Models
by: Huang, Jiabo, et al.
Published: (2026)
by: Huang, Jiabo, et al.
Published: (2026)
Seeing Further on the Shoulders of Giants: Knowledge Inheritance for Vision Foundation Models
by: Huang, Jiabo, et al.
Published: (2025)
by: Huang, Jiabo, et al.
Published: (2025)
Evaluating Fairness in Large Vision-Language Models Across Diverse Demographic Attributes and Prompts
by: Wu, Xuyang, et al.
Published: (2024)
by: Wu, Xuyang, et al.
Published: (2024)
Compact Vision Transformer by Reduction of Kernel Complexity
by: Wang, Yancheng, et al.
Published: (2025)
by: Wang, Yancheng, et al.
Published: (2025)
CoFie: Learning Compact Neural Surface Representations with Coordinate Fields
by: Jiang, Hanwen, et al.
Published: (2024)
by: Jiang, Hanwen, et al.
Published: (2024)
Local-Global Context Aware Transformer for Language-Guided Video Segmentation
by: Liang, Chen, et al.
Published: (2022)
by: Liang, Chen, et al.
Published: (2022)
Sparse-Tuning: Adapting Vision Transformers with Efficient Fine-tuning and Inference
by: Liu, Ting, et al.
Published: (2024)
by: Liu, Ting, et al.
Published: (2024)
Robust Scene Coordinate Regression via Geometrically-Consistent Global Descriptors
by: Nguyen, Son Tung, et al.
Published: (2025)
by: Nguyen, Son Tung, et al.
Published: (2025)
Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models
by: Liu, Xuyang, et al.
Published: (2025)
by: Liu, Xuyang, et al.
Published: (2025)
GlobalMamba: Global Image Serialization for Vision Mamba
by: Wang, Chengkun, et al.
Published: (2024)
by: Wang, Chengkun, et al.
Published: (2024)
Unleashing Vision Transformer Potential In Image Quality Assessment via Global-Local Adaptive Interaction
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
by: Wang, Ziyao, et al.
Published: (2026)
by: Wang, Ziyao, et al.
Published: (2026)
Robust Partial 3D Point Cloud Registration via Confidence Estimation under Global Context
by: Wang, Yongqiang, et al.
Published: (2025)
by: Wang, Yongqiang, et al.
Published: (2025)
Variation-aware Vision Token Dropping for Faster Large Vision-Language Models
by: Chen, Junjie, et al.
Published: (2025)
by: Chen, Junjie, et al.
Published: (2025)
Token Compression Meets Compact Vision Transformers: A Survey and Comparative Evaluation for Edge AI
by: Nguyen, Phat, et al.
Published: (2025)
by: Nguyen, Phat, et al.
Published: (2025)
Multi-Attribute Vision Transformers are Efficient and Robust Learners
by: Gani, Hanan, et al.
Published: (2024)
by: Gani, Hanan, et al.
Published: (2024)
Distilling Vision Transformers for Distortion-Robust Representation Learning
by: Alexis, Konstantinos, et al.
Published: (2026)
by: Alexis, Konstantinos, et al.
Published: (2026)
Global Occlusion-Aware Transformer for Robust Stereo Matching
by: Liu, Zihua, et al.
Published: (2023)
by: Liu, Zihua, et al.
Published: (2023)
LSM-YOLO: A Compact and Effective ROI Detector for Medical Detection
by: Yu, Zhongwen, et al.
Published: (2024)
by: Yu, Zhongwen, et al.
Published: (2024)
Learning Occlusion-Robust Vision Transformers for Real-Time UAV Tracking
by: Wu, You, et al.
Published: (2025)
by: Wu, You, et al.
Published: (2025)
hZACH-ViT: Curved Latent Geometry for Compact Vision Transformers in Low-Data Medical Imaging
by: Angelakis, Athanasios
Published: (2026)
by: Angelakis, Athanasios
Published: (2026)
In the Eye of Transformer: Global-Local Correlation for Egocentric Gaze Estimation
by: Lai, Bolin, et al.
Published: (2022)
by: Lai, Bolin, et al.
Published: (2022)
GLACE: Global Local Accelerated Coordinate Encoding
by: Wang, Fangjinhua, et al.
Published: (2024)
by: Wang, Fangjinhua, et al.
Published: (2024)
On the Limits of Token Reduction for Efficient Unified Vision Language Training
by: Chen, Siyi, et al.
Published: (2026)
by: Chen, Siyi, et al.
Published: (2026)
Global Average Feature Augmentation for Robust Semantic Segmentation with Transformers
by: Salgado, Alberto Gonzalo Rodriguez, et al.
Published: (2024)
by: Salgado, Alberto Gonzalo Rodriguez, et al.
Published: (2024)
Learning Motion Blur Robust Vision Transformers for Real-Time UAV Tracking
by: Wu, You, et al.
Published: (2024)
by: Wu, You, et al.
Published: (2024)
Dense Vision Transformer Compression with Few Samples
by: Zhang, Hanxiao, et al.
Published: (2024)
by: Zhang, Hanxiao, et al.
Published: (2024)
ESsEN: Training Compact Discriminative Vision-Language Transformers in a Low-Resource Setting
by: Fields, Clayton, et al.
Published: (2026)
by: Fields, Clayton, et al.
Published: (2026)
Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models
by: Liu, Xuyang, et al.
Published: (2025)
by: Liu, Xuyang, et al.
Published: (2025)
ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions
by: Xia, Chunlong, et al.
Published: (2024)
by: Xia, Chunlong, et al.
Published: (2024)
Vision Transformer for Robust Occluded Person Reidentification in Complex Surveillance Scenes
by: Li, Bo, et al.
Published: (2025)
by: Li, Bo, et al.
Published: (2025)
Patch-Fool: Are Vision Transformers Always Robust Against Adversarial Perturbations?
by: Fu, Yonggan, et al.
Published: (2022)
by: Fu, Yonggan, et al.
Published: (2022)
FullLoRA: Efficiently Boosting the Robustness of Pretrained Vision Transformers
by: Yuan, Zheng, et al.
Published: (2024)
by: Yuan, Zheng, et al.
Published: (2024)
Approximate Nullspace Augmented Finetuning for Robust Vision Transformers
by: Liu, Haoyang, et al.
Published: (2024)
by: Liu, Haoyang, et al.
Published: (2024)
When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models
by: Yan, Yuping, et al.
Published: (2025)
by: Yan, Yuping, et al.
Published: (2025)
Similar Items
-
CAViT -- Channel-Aware Vision Transformer for Dynamic Feature Fusion
by: Safdar, Aon, et al.
Published: (2026) -
Faster and Better: Reinforced Collaborative Distillation and Self-Learning for Infrared-Visible Image Fusion
by: Wang, Yuhao, et al.
Published: (2025) -
Infrared and Visible Image Fusion with Language-Driven Loss in CLIP Embedding Space
by: Wang, Yuhao, et al.
Published: (2024) -
Semantic Object-level Modeling for Robust Visual Camera Relocalization
by: Zhu, Yifan, et al.
Published: (2024) -
CAViAR: Critic-Augmented Video Agentic Reasoning
by: Menon, Sachit, et al.
Published: (2025)