SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement
Fuente:
arXiv
Saved in:
| Main Authors: | Yin, Weijie, Yang, Dingkang, Dong, Hongyuan, Kang, Zijian, Wang, Jiacong, Liang, Xiao, Feng, Chao, Ran, Jiao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scalable Vision Language Model Training via High Quality Data Curation
by: Dong, Hongyuan, et al.
Published: (2025)
by: Dong, Hongyuan, et al.
Published: (2025)
AdaLRS: Loss-Guided Adaptive Learning Rate Search for Efficient Foundation Model Pretraining
by: Dong, Hongyuan, et al.
Published: (2025)
by: Dong, Hongyuan, et al.
Published: (2025)
SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
by: Shu, Fangxun, et al.
Published: (2025)
by: Shu, Fangxun, et al.
Published: (2025)
SAIL-VL2 Technical Report
by: Yin, Weijie, et al.
Published: (2025)
by: Yin, Weijie, et al.
Published: (2025)
VGR: Visual Grounded Reasoning
by: Wang, Jiacong, et al.
Published: (2025)
by: Wang, Jiacong, et al.
Published: (2025)
Multifaceted Evaluation of Audio-Visual Capability for MLLMs: Effectiveness, Efficiency, Generalizability and Robustness
by: Zhao, Yusheng, et al.
Published: (2025)
by: Zhao, Yusheng, et al.
Published: (2025)
Learning, Reasoning, Refinement: A Framework for Kahneman's Dual-System Intelligence in GUI Agents
by: Wei, Jinjie, et al.
Published: (2025)
by: Wei, Jinjie, et al.
Published: (2025)
Explore How to Inject Beneficial Noise in MLLMs
by: Zhu, Ruishu, et al.
Published: (2025)
by: Zhu, Ruishu, et al.
Published: (2025)
Towards Generalizable Robotic Manipulation in Dynamic Environments
by: Fang, Heng, et al.
Published: (2026)
by: Fang, Heng, et al.
Published: (2026)
BreathNet: Generalizable Audio Deepfake Detection via Breath-Cue-Guided Feature Refinement
by: Ye, Zhe, et al.
Published: (2026)
by: Ye, Zhe, et al.
Published: (2026)
Towards Faithful Reasoning in Comics for Small MLLMs
by: Feng, Chengcheng, et al.
Published: (2026)
by: Feng, Chengcheng, et al.
Published: (2026)
Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
by: Xiao, Kelaiti, et al.
Published: (2025)
by: Xiao, Kelaiti, et al.
Published: (2025)
Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs
by: Jiang, Xueying, et al.
Published: (2026)
by: Jiang, Xueying, et al.
Published: (2026)
On the Importance of Backbone to the Adversarial Robustness of Object Detectors
by: Li, Xiao, et al.
Published: (2023)
by: Li, Xiao, et al.
Published: (2023)
UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?
by: Liu, Yuanxin, et al.
Published: (2025)
by: Liu, Yuanxin, et al.
Published: (2025)
GSsplat: Generalizable Semantic Gaussian Splatting for Novel-view Synthesis in 3D Scenes
by: Xiao, Feng, et al.
Published: (2025)
by: Xiao, Feng, et al.
Published: (2025)
Benchmarking and Improving Detail Image Caption
by: Dong, Hongyuan, et al.
Published: (2024)
by: Dong, Hongyuan, et al.
Published: (2024)
Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models
by: Ma, Qihang, et al.
Published: (2025)
by: Ma, Qihang, et al.
Published: (2025)
Improving Adversarial Robustness via Feature Pattern Consistency Constraint
by: Hu, Jiacong, et al.
Published: (2024)
by: Hu, Jiacong, et al.
Published: (2024)
Towards Robust Evaluation of STEM Education: Leveraging MLLMs in Project-Based Learning
by: Wu, Xinyi, et al.
Published: (2025)
by: Wu, Xinyi, et al.
Published: (2025)
ViC-Bench: Benchmarking Visual-Interleaved Chain-of-Thought Capability in MLLMs with Free-Style Intermediate State Representations
by: Wu, Xuecheng, et al.
Published: (2025)
by: Wu, Xuecheng, et al.
Published: (2025)
Gradual Metaprogramming
by: Chen, Tianyu, et al.
Published: (2025)
by: Chen, Tianyu, et al.
Published: (2025)
IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
by: Jiang, Yankai, et al.
Published: (2026)
by: Jiang, Yankai, et al.
Published: (2026)
QG-VTC: Question-Guided Visual Token Compression in MLLMs for Efficient VQA
by: Li, Shuai, et al.
Published: (2025)
by: Li, Shuai, et al.
Published: (2025)
SecCoder: Towards Generalizable and Robust Secure Code Generation
by: Zhang, Boyu, et al.
Published: (2024)
by: Zhang, Boyu, et al.
Published: (2024)
PRVQL: Progressive Knowledge-guided Refinement for Robust Egocentric Visual Query Localization
by: Fan, Bing, et al.
Published: (2025)
by: Fan, Bing, et al.
Published: (2025)
ReVis: Towards Reusable Image-Based Visualizations with MLLMs
by: Wen, Xiaolin, et al.
Published: (2026)
by: Wen, Xiaolin, et al.
Published: (2026)
DGIQA: Depth-guided Feature Attention and Refinement for Generalizable Image Quality Assessment
by: Ramesh, Vaishnav, et al.
Published: (2025)
by: Ramesh, Vaishnav, et al.
Published: (2025)
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
by: Tang, Yolo Y., et al.
Published: (2025)
by: Tang, Yolo Y., et al.
Published: (2025)
Torsion-Space Diffusion for Protein Backbone Generation with Geometric Refinement
by: Singh, Lakshaditya, et al.
Published: (2025)
by: Singh, Lakshaditya, et al.
Published: (2025)
Generalizable Image Repair for Robust Visual Control
by: Sobolewski, Carson, et al.
Published: (2025)
by: Sobolewski, Carson, et al.
Published: (2025)
Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones
by: Teleki, Maria, et al.
Published: (2025)
by: Teleki, Maria, et al.
Published: (2025)
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities
by: Jiang, Shixin, et al.
Published: (2024)
by: Jiang, Shixin, et al.
Published: (2024)
Iterative Flow Matching -- Path Correction and Gradual Refinement for Enhanced Generative Modeling
by: Haber, Eldad, et al.
Published: (2025)
by: Haber, Eldad, et al.
Published: (2025)
RL from Teacher-Model Refinement: Gradual Imitation Learning for Machine Translation
by: Lee, Dongyub Jude, et al.
Published: (2025)
by: Lee, Dongyub Jude, et al.
Published: (2025)
SIFT-Graph: Benchmarking Multimodal Defense Against Image Adversarial Attacks With Robust Feature Graph
by: He, Jingjie, et al.
Published: (2025)
by: He, Jingjie, et al.
Published: (2025)
CodePercept: Code-Grounded Visual STEM Perception for MLLMs
by: Guan, Tongkun, et al.
Published: (2026)
by: Guan, Tongkun, et al.
Published: (2026)
Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
AdaCodec: A Predictive Visual Code for Video MLLMs
by: Hou, Haowen, et al.
Published: (2026)
by: Hou, Haowen, et al.
Published: (2026)
Similar Items
-
Scalable Vision Language Model Training via High Quality Data Curation
by: Dong, Hongyuan, et al.
Published: (2025) -
AdaLRS: Loss-Guided Adaptive Learning Rate Search for Efficient Foundation Model Pretraining
by: Dong, Hongyuan, et al.
Published: (2025) -
SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
by: Shu, Fangxun, et al.
Published: (2025) -
SAIL-VL2 Technical Report
by: Yin, Weijie, et al.
Published: (2025) -
VGR: Visual Grounded Reasoning
by: Wang, Jiacong, et al.
Published: (2025)