ViCO: A Training Strategy towards Semantic Aware Dynamic High-Resolution
Fuente:
arXiv
Saved in:
| Main Authors: | Cui, Long, Wang, Weiyun, Shao, Jie, Wen, Zichen, Luo, Gen, Zhang, Linfeng, Zhang, Yanting, Qiao, Yu, Wang, Wenhai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models
by: Yang, Yantai, et al.
Published: (2025)
by: Yang, Yantai, et al.
Published: (2025)
Flash-Unified: A Training-Free and Task-Aware Acceleration Framework for Native Unified Models
by: Ke, Junlong, et al.
Published: (2026)
by: Ke, Junlong, et al.
Published: (2026)
Sequential Diffusion Language Models
by: Liu, Yangzhou, et al.
Published: (2025)
by: Liu, Yangzhou, et al.
Published: (2025)
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
by: Luo, Gen, et al.
Published: (2025)
by: Luo, Gen, et al.
Published: (2025)
UniViTAR: Unified Vision Transformer with Native Resolution
by: Qiao, Limeng, et al.
Published: (2025)
by: Qiao, Limeng, et al.
Published: (2025)
AgentSlimming: Towards Efficient and Cost-Aware Multi-Agent Systems
by: Chen, Yulang, et al.
Published: (2026)
by: Chen, Yulang, et al.
Published: (2026)
NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
by: Tian, Changyao, et al.
Published: (2025)
by: Tian, Changyao, et al.
Published: (2025)
Point or Line? Using Line-based Representation for Panoptic Symbol Spotting in CAD Drawings
by: Wei, Xingguang, et al.
Published: (2025)
by: Wei, Xingguang, et al.
Published: (2025)
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
by: Lei, Zhenxin, et al.
Published: (2025)
by: Lei, Zhenxin, et al.
Published: (2025)
OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis
by: Chen, Junting, et al.
Published: (2025)
by: Chen, Junting, et al.
Published: (2025)
GenExam: A Multidisciplinary Text-to-Image Exam
by: Wang, Zhaokai, et al.
Published: (2025)
by: Wang, Zhaokai, et al.
Published: (2025)
JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search
by: Zou, Dongyun, et al.
Published: (2026)
by: Zou, Dongyun, et al.
Published: (2026)
Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning
by: Ke, Junlong, et al.
Published: (2026)
by: Ke, Junlong, et al.
Published: (2026)
Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?
by: Wen, Zichen, et al.
Published: (2025)
by: Wen, Zichen, et al.
Published: (2025)
Outlier-Aware Post-Training Quantization for Image Super-Resolution
by: Wang, Hailing, et al.
Published: (2025)
by: Wang, Hailing, et al.
Published: (2025)
Dynamic-Aware Video Distillation: Optimizing Temporal Resolution Based on Video Semantics
by: Zhao, Yinjie, et al.
Published: (2025)
by: Zhao, Yinjie, et al.
Published: (2025)
Diffusion LLM with Native Variable Generation Lengths: Let [EOS] Lead the Way
by: Yang, Yicun, et al.
Published: (2025)
by: Yang, Yicun, et al.
Published: (2025)
Learning Optimal Distributionally Robust Individualized Treatment Rules Integrating Multi-Source Data
by: Cui, Wenhai, et al.
Published: (2026)
by: Cui, Wenhai, et al.
Published: (2026)
STAMICS: Splat, Track And Map with Integrated Consistency and Semantics for Dense RGB-D SLAM
by: Wang, Yongxu, et al.
Published: (2025)
by: Wang, Yongxu, et al.
Published: (2025)
Demystify Transformers & Convolutions in Modern Image Deep Networks
by: Hu, Xiaowei, et al.
Published: (2022)
by: Hu, Xiaowei, et al.
Published: (2022)
Semantic-Aware Adversarial Training for Reliable Deep Hashing Retrieval
by: Yuan, Xu, et al.
Published: (2023)
by: Yuan, Xu, et al.
Published: (2023)
Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
by: Wen, Zichen, et al.
Published: (2025)
by: Wen, Zichen, et al.
Published: (2025)
Needle In A Multimodal Haystack
by: Wang, Weiyun, et al.
Published: (2024)
by: Wang, Weiyun, et al.
Published: (2024)
Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
by: Duan, Yuchen, et al.
Published: (2024)
by: Duan, Yuchen, et al.
Published: (2024)
A Fault Diagnosis Method for Avionics Equipment Based on SMOTEWB‐LGBM
by: Gen Li, et al.
Published: (2024)
by: Gen Li, et al.
Published: (2024)
ExtraVAR: Stage-Aware RoPE Remapping for Resolution Extrapolation in Visual Autoregressive Models
by: Yan, Feihong, et al.
Published: (2026)
by: Yan, Feihong, et al.
Published: (2026)
Trio-ViT: Post-Training Quantization and Acceleration for Softmax-Free Efficient Vision Transformer
by: Shi, Huihong, et al.
Published: (2024)
by: Shi, Huihong, et al.
Published: (2024)
Thinking Inside the Mask: In-Place Prompting in Diffusion LLMs
by: Jin, Xiangqi, et al.
Published: (2025)
by: Jin, Xiangqi, et al.
Published: (2025)
OminiAdapt: Learning Cross-Task Invariance for Robust and Environment-Aware Robotic Manipulation
by: Wang, Yongxu, et al.
Published: (2025)
by: Wang, Yongxu, et al.
Published: (2025)
Maximizing the spectral radius of graphs of given size with forbidden a subgraph
by: Zhang, Yanting, et al.
Published: (2024)
by: Zhang, Yanting, et al.
Published: (2024)
The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
by: Wang, Weiyun, et al.
Published: (2024)
by: Wang, Weiyun, et al.
Published: (2024)
REEF: Representation Encoding Fingerprints for Large Language Models
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
by: Yang, Ganlin, et al.
Published: (2025)
by: Yang, Ganlin, et al.
Published: (2025)
SiamSeg: Self-Training with Contrastive Learning for Unsupervised Domain Adaptation Semantic Segmentation in Remote Sensing
by: Wang, Bin, et al.
Published: (2024)
by: Wang, Bin, et al.
Published: (2024)
Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion
by: Yi, Xunpeng, et al.
Published: (2024)
by: Yi, Xunpeng, et al.
Published: (2024)
Harnessing Ash for Sustainable CO2 Absorption: Current Strategies and Future Prospects
by: Wen‐Ya Wu, et al.
Published: (2024)
by: Wen‐Ya Wu, et al.
Published: (2024)
PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding
by: Zhang, Haoze, et al.
Published: (2025)
by: Zhang, Haoze, et al.
Published: (2025)
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
by: Chen, Mengzhao, et al.
Published: (2024)
by: Chen, Mengzhao, et al.
Published: (2024)
ViPOcc: Leveraging Visual Priors from Vision Foundation Models for Single-View 3D Occupancy Prediction
by: Feng, Yi, et al.
Published: (2024)
by: Feng, Yi, et al.
Published: (2024)
ZO-SAM: Zero-Order Sharpness-Aware Minimization for Efficient Sparse Training
by: Ji, Jie, et al.
Published: (2026)
by: Ji, Jie, et al.
Published: (2026)
Similar Items
-
EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models
by: Yang, Yantai, et al.
Published: (2025) -
Flash-Unified: A Training-Free and Task-Aware Acceleration Framework for Native Unified Models
by: Ke, Junlong, et al.
Published: (2026) -
Sequential Diffusion Language Models
by: Liu, Yangzhou, et al.
Published: (2025) -
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
by: Luo, Gen, et al.
Published: (2025) -
UniViTAR: Unified Vision Transformer with Native Resolution
by: Qiao, Limeng, et al.
Published: (2025)