Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Simon, Panigrahi, Abhishek, Cheng, Yun, Yu, Dingli, Goyal, Anirudh, Arora, Sanjeev |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Models Learn Skill Composition from Examples?
by: Zhao, Haoyu, et al.
Published: (2024)
by: Zhao, Haoyu, et al.
Published: (2024)
Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning
by: Kaur, Simran, et al.
Published: (2024)
by: Kaur, Simran, et al.
Published: (2024)
Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates
by: Lyu, Kaifeng, et al.
Published: (2024)
by: Lyu, Kaifeng, et al.
Published: (2024)
AdaptMI: Adaptive Skill-based In-context Math Instruction for Small Language Models
by: He, Yinghui, et al.
Published: (2025)
by: He, Yinghui, et al.
Published: (2025)
Trainable Transformer in Transformer
by: Panigrahi, Abhishek, et al.
Published: (2023)
by: Panigrahi, Abhishek, et al.
Published: (2023)
Contextual Drag: How Errors in the Context Affect LLM Reasoning
by: Cheng, Yun, et al.
Published: (2026)
by: Cheng, Yun, et al.
Published: (2026)
Escape the Language Prior: Mitigating Late-Stage Modality Collapse in Audio Reasoning via Modality-Aware Policy Optimization
by: Xiao, Cihan, et al.
Published: (2026)
by: Xiao, Cihan, et al.
Published: (2026)
Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs
by: Zhu, Mengdan, et al.
Published: (2026)
by: Zhu, Mengdan, et al.
Published: (2026)
Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
by: He, Yinghui, et al.
Published: (2026)
by: He, Yinghui, et al.
Published: (2026)
Can VLMs Recall Factual Associations From Visual References?
by: Ashok, Dhananjay, et al.
Published: (2025)
by: Ashok, Dhananjay, et al.
Published: (2025)
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
by: Mayer, Julius, et al.
Published: (2025)
by: Mayer, Julius, et al.
Published: (2025)
Can We Trust LLMs on Memristors? Diving into Reasoning Ability under Non-Ideality
by: Wu, Taiqiang, et al.
Published: (2026)
by: Wu, Taiqiang, et al.
Published: (2026)
ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty
by: Wu, Xindi, et al.
Published: (2024)
by: Wu, Xindi, et al.
Published: (2024)
Beyond Logit Lens: Contextual Embeddings for Robust Hallucination Detection & Grounding in VLMs
by: Phukan, Anirudh, et al.
Published: (2024)
by: Phukan, Anirudh, et al.
Published: (2024)
LinkTransformer: A Unified Package for Record Linkage with Transformer Language Models
by: Arora, Abhishek, et al.
Published: (2023)
by: Arora, Abhishek, et al.
Published: (2023)
Can We Trust LLMs? Mitigate Overconfidence Bias in LLMs through Knowledge Transfer
by: Yang, Haoyan, et al.
Published: (2024)
by: Yang, Haoyan, et al.
Published: (2024)
Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
by: Pan, Zhiyu, et al.
Published: (2026)
by: Pan, Zhiyu, et al.
Published: (2026)
LLMs Can Evolve Continually on Modality for X-Modal Reasoning
by: Yu, Jiazuo, et al.
Published: (2024)
by: Yu, Jiazuo, et al.
Published: (2024)
ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context
by: Kim, Joongwon, et al.
Published: (2025)
by: Kim, Joongwon, et al.
Published: (2025)
Reasoning Robustness of LLMs to Adversarial Typographical Errors
by: Gan, Esther, et al.
Published: (2024)
by: Gan, Esther, et al.
Published: (2024)
Do Multilingual VLMs Reason Equally? A Cross-Lingual Visual Reasoning Audit for Indian Languages
by: R, Swastik
Published: (2026)
by: R, Swastik
Published: (2026)
[De|Re]constructing VLMs' Reasoning in Counting
by: Alghisi, Simone, et al.
Published: (2025)
by: Alghisi, Simone, et al.
Published: (2025)
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
by: Shi, Chufan, et al.
Published: (2026)
by: Shi, Chufan, et al.
Published: (2026)
When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs
by: Penamakuri, Abhirama Subramanyam, et al.
Published: (2025)
by: Penamakuri, Abhirama Subramanyam, et al.
Published: (2025)
Can We Trust LLM Detectors?
by: Sandhan, Jivnesh, et al.
Published: (2026)
by: Sandhan, Jivnesh, et al.
Published: (2026)
Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs
by: Carbune, Victor, et al.
Published: (2024)
by: Carbune, Victor, et al.
Published: (2024)
Do LLMs and VLMs Share Neurons for Inference? Evidence and Mechanisms of Cross-Modal Transfer
by: Cui, Chenhang, et al.
Published: (2026)
by: Cui, Chenhang, et al.
Published: (2026)
MMGR: Multi-Modal Generative Reasoning
by: Cai, Zefan, et al.
Published: (2025)
by: Cai, Zefan, et al.
Published: (2025)
Representing Rule-based Chatbots with Transformers
by: Friedman, Dan, et al.
Published: (2024)
by: Friedman, Dan, et al.
Published: (2024)
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
by: Wang, Dianyi, et al.
Published: (2025)
by: Wang, Dianyi, et al.
Published: (2025)
Mitigating Data Imbalance in Automated Speaking Assessment
by: Tsai, Fong-Chun, et al.
Published: (2025)
by: Tsai, Fong-Chun, et al.
Published: (2025)
VLMs Can Aggregate Scattered Training Patches
by: Zhou, Zhanhui, et al.
Published: (2025)
by: Zhou, Zhanhui, et al.
Published: (2025)
Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
by: Bharadwaj, Anirudh, et al.
Published: (2025)
by: Bharadwaj, Anirudh, et al.
Published: (2025)
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
by: Shen, Yifan, et al.
Published: (2025)
by: Shen, Yifan, et al.
Published: (2025)
Improving Sparse Memory Finetuning
by: Goyal, Satyam, et al.
Published: (2026)
by: Goyal, Satyam, et al.
Published: (2026)
Sparse Memory Finetuning as a Low-Forgetting Alternative to LoRA and Full Finetuning
by: Gupta, Prakhar, et al.
Published: (2026)
by: Gupta, Prakhar, et al.
Published: (2026)
Accelerating Greedy Coordinate Gradient and General Prompt Optimization via Probe Sampling
by: Zhao, Yiran, et al.
Published: (2024)
by: Zhao, Yiran, et al.
Published: (2024)
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
by: Gao, Xin, et al.
Published: (2026)
by: Gao, Xin, et al.
Published: (2026)
How Far Are We from Intelligent Visual Deductive Reasoning?
by: Zhang, Yizhe, et al.
Published: (2024)
by: Zhang, Yizhe, et al.
Published: (2024)
Comparing Importance Sampling Based Methods for Mitigating the Effect of Class Imbalance
by: Panigrahi, Indu, et al.
Published: (2024)
by: Panigrahi, Indu, et al.
Published: (2024)
Similar Items
-
Can Models Learn Skill Composition from Examples?
by: Zhao, Haoyu, et al.
Published: (2024) -
Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning
by: Kaur, Simran, et al.
Published: (2024) -
Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates
by: Lyu, Kaifeng, et al.
Published: (2024) -
AdaptMI: Adaptive Skill-based In-context Math Instruction for Small Language Models
by: He, Yinghui, et al.
Published: (2025) -
Trainable Transformer in Transformer
by: Panigrahi, Abhishek, et al.
Published: (2023)