Better Reasoning with Less Data: Enhancing VLMs Through Unified Modality Scoring
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Mingjie, Estornell, Andrew, Yang, Hongzheng, Zhao, Yuzhi, Zhu, Zhaowei, Xuan, Qi, Wei, Jiaheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evian: Towards Explainable Visual Instruction-tuning Data Auditing
von: Jia, Zimu, et al.
Veröffentlicht: (2026)
von: Jia, Zimu, et al.
Veröffentlicht: (2026)
AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
von: Liu, Xianyang, et al.
Veröffentlicht: (2025)
von: Liu, Xianyang, et al.
Veröffentlicht: (2025)
SelectMix: Enhancing Label Noise Robustness through Targeted Sample Mixing
von: Liu, Qiuhao, et al.
Veröffentlicht: (2025)
von: Liu, Qiuhao, et al.
Veröffentlicht: (2025)
GRAID: Enhancing Spatial Reasoning of VLMs Through High-Fidelity Data Generation
von: Elmaaroufi, Karim, et al.
Veröffentlicht: (2025)
von: Elmaaroufi, Karim, et al.
Veröffentlicht: (2025)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
von: Zhang, Jianrui, et al.
Veröffentlicht: (2026)
von: Zhang, Jianrui, et al.
Veröffentlicht: (2026)
Reassessing Layer Pruning in LLMs: New Insights and Methods
von: Lu, Yao, et al.
Veröffentlicht: (2024)
von: Lu, Yao, et al.
Veröffentlicht: (2024)
LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations
von: Xu, Mingjie, et al.
Veröffentlicht: (2024)
von: Xu, Mingjie, et al.
Veröffentlicht: (2024)
VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models
von: Xu, Mingjie, et al.
Veröffentlicht: (2025)
von: Xu, Mingjie, et al.
Veröffentlicht: (2025)
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
von: Huang, Ting, et al.
Veröffentlicht: (2025)
von: Huang, Ting, et al.
Veröffentlicht: (2025)
Activating Visual Context and Commonsense Reasoning through Masked Prediction in VLMs
von: Yu, Jiaao, et al.
Veröffentlicht: (2025)
von: Yu, Jiaao, et al.
Veröffentlicht: (2025)
Vision Verification Enhanced Fusion of VLMs for Efficient Visual Reasoning
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2026)
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2026)
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
von: Wang, Dianyi, et al.
Veröffentlicht: (2025)
von: Wang, Dianyi, et al.
Veröffentlicht: (2025)
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
von: Clark, Christopher, et al.
Veröffentlicht: (2026)
von: Clark, Christopher, et al.
Veröffentlicht: (2026)
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
von: Zhu, Boyu, et al.
Veröffentlicht: (2025)
von: Zhu, Boyu, et al.
Veröffentlicht: (2025)
From Shots to Stories: LLM-Assisted Video Editing with Unified Language Representations
von: Li, Yuzhi, et al.
Veröffentlicht: (2025)
von: Li, Yuzhi, et al.
Veröffentlicht: (2025)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
von: Ge, Yuyao, et al.
Veröffentlicht: (2025)
von: Ge, Yuyao, et al.
Veröffentlicht: (2025)
CLGRPO: Reasoning Ability Enhancement for Small VLMs
von: Wang, Fanyi, et al.
Veröffentlicht: (2025)
von: Wang, Fanyi, et al.
Veröffentlicht: (2025)
TAR-TVG: Enhancing VLMs with Timestamp Anchor-Constrained Reasoning for Temporal Video Grounding
von: Guo, Chaohong, et al.
Veröffentlicht: (2025)
von: Guo, Chaohong, et al.
Veröffentlicht: (2025)
Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
von: Pan, Zhiyu, et al.
Veröffentlicht: (2026)
von: Pan, Zhiyu, et al.
Veröffentlicht: (2026)
Unlocking Dense Metric Depth Estimation in VLMs
von: Yu, Hanxun, et al.
Veröffentlicht: (2026)
von: Yu, Hanxun, et al.
Veröffentlicht: (2026)
Better with Less: Tackling Heterogeneous Multi-Modal Image Joint Pretraining via Conditioned and Degraded Masked Autoencoder
von: Peng, Bowen, et al.
Veröffentlicht: (2026)
von: Peng, Bowen, et al.
Veröffentlicht: (2026)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
von: Qin, Yiming, et al.
Veröffentlicht: (2025)
von: Qin, Yiming, et al.
Veröffentlicht: (2025)
Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
von: Lu, Meng, et al.
Veröffentlicht: (2025)
von: Lu, Meng, et al.
Veröffentlicht: (2025)
Enhancing Multimodal Unified Representations for Cross Modal Generalization
von: Huang, Hai, et al.
Veröffentlicht: (2024)
von: Huang, Hai, et al.
Veröffentlicht: (2024)
Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards
von: Gambashidze, Alexander, et al.
Veröffentlicht: (2025)
von: Gambashidze, Alexander, et al.
Veröffentlicht: (2025)
Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement Learning
von: Li, Yue, et al.
Veröffentlicht: (2025)
von: Li, Yue, et al.
Veröffentlicht: (2025)
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering
von: Jia, Yiduo, et al.
Veröffentlicht: (2026)
von: Jia, Yiduo, et al.
Veröffentlicht: (2026)
AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
von: Zhang, Junyang, et al.
Veröffentlicht: (2025)
von: Zhang, Junyang, et al.
Veröffentlicht: (2025)
ABC: Achieving Better Control of Multimodal Embeddings using VLMs
von: Schneider, Benjamin, et al.
Veröffentlicht: (2025)
von: Schneider, Benjamin, et al.
Veröffentlicht: (2025)
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
von: Shen, Yifan, et al.
Veröffentlicht: (2025)
von: Shen, Yifan, et al.
Veröffentlicht: (2025)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
von: Li, Weiming, et al.
Veröffentlicht: (2025)
von: Li, Weiming, et al.
Veröffentlicht: (2025)
BetterCheck: Towards Safeguarding VLMs for Automotive Perception Systems
von: Dona, Malsha Ashani Mahawatta, et al.
Veröffentlicht: (2025)
von: Dona, Malsha Ashani Mahawatta, et al.
Veröffentlicht: (2025)
Data Factory with Minimal Human Effort Using VLMs
von: Ye, Jiaojiao, et al.
Veröffentlicht: (2025)
von: Ye, Jiaojiao, et al.
Veröffentlicht: (2025)
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
von: Park, Simon, et al.
Veröffentlicht: (2025)
von: Park, Simon, et al.
Veröffentlicht: (2025)
SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs
von: Bo, Zi-Hao, et al.
Veröffentlicht: (2026)
von: Bo, Zi-Hao, et al.
Veröffentlicht: (2026)
Training-Free Dual Hyperbolic Adapters for Better Cross-Modal Reasoning
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
von: Berman, Shmuel, et al.
Veröffentlicht: (2025)
von: Berman, Shmuel, et al.
Veröffentlicht: (2025)
3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering
von: Xu, Rongtao, et al.
Veröffentlicht: (2025)
von: Xu, Rongtao, et al.
Veröffentlicht: (2025)
Caption This, Reason That: VLMs Caught in the Middle
von: Weng, Zihan, et al.
Veröffentlicht: (2025)
von: Weng, Zihan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Evian: Towards Explainable Visual Instruction-tuning Data Auditing
von: Jia, Zimu, et al.
Veröffentlicht: (2026) -
AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
von: Liu, Xianyang, et al.
Veröffentlicht: (2025) -
SelectMix: Enhancing Label Noise Robustness through Targeted Sample Mixing
von: Liu, Qiuhao, et al.
Veröffentlicht: (2025) -
GRAID: Enhancing Spatial Reasoning of VLMs Through High-Fidelity Data Generation
von: Elmaaroufi, Karim, et al.
Veröffentlicht: (2025) -
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
von: Zhang, Jianrui, et al.
Veröffentlicht: (2026)