More Than the Final Answer: Improving Visual Extraction and Logical Consistency in Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Just, Hoang Anh, Fan, Yifei, Zhao, Handong, Gu, Jiuxiang, Zhang, Ruiyi, Jenni, Simon, Kafle, Kushal, Jia, Ruoxi, Shi, Jing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Seeing Through Words: Controlling Visual Retrieval Quality with Language Models
by: Lu, Jianglin, et al.
Published: (2026)
by: Lu, Jianglin, et al.
Published: (2026)
Improving Large Vision and Language Models by Learning from a Panel of Peers
by: Hernandez, Jefferson, et al.
Published: (2025)
by: Hernandez, Jefferson, et al.
Published: (2025)
MAGNET: Augmenting Generative Decoders with Representation Learning and Infilling Capabilities
by: Khosla, Savya, et al.
Published: (2025)
by: Khosla, Savya, et al.
Published: (2025)
RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward
by: Wu, Qiucheng, et al.
Published: (2026)
by: Wu, Qiucheng, et al.
Published: (2026)
Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
by: Yang, Ziyan, et al.
Published: (2022)
by: Yang, Ziyan, et al.
Published: (2022)
The Signal is in the Steps: Local Scoring for Reasoning Data Selection
by: Just, Hoang Anh, et al.
Published: (2025)
by: Just, Hoang Anh, et al.
Published: (2025)
Data-Centric Human Preference with Rationales for Direct Preference Alignment
by: Just, Hoang Anh, et al.
Published: (2024)
by: Just, Hoang Anh, et al.
Published: (2024)
Probing Knowledge Holes in Unlearned LLMs
by: Ko, Myeongseob, et al.
Published: (2025)
by: Ko, Myeongseob, et al.
Published: (2025)
DiPT: Enhancing LLM reasoning through diversified perspective-taking
by: Just, Hoang Anh, et al.
Published: (2024)
by: Just, Hoang Anh, et al.
Published: (2024)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
by: Shrestha, Robik, et al.
Published: (2020)
by: Shrestha, Robik, et al.
Published: (2020)
CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
by: Dong, Qihua, et al.
Published: (2025)
by: Dong, Qihua, et al.
Published: (2025)
FINEMATCH: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction
by: Hua, Hang, et al.
Published: (2024)
by: Hua, Hang, et al.
Published: (2024)
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
by: Zang, Yuan, et al.
Published: (2025)
by: Zang, Yuan, et al.
Published: (2025)
Customization Assistant for Text-to-image Generation
by: Zhou, Yufan, et al.
Published: (2023)
by: Zhou, Yufan, et al.
Published: (2023)
Plot'n Polish: Zero-shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models
by: Akdemir, Kiymet, et al.
Published: (2025)
by: Akdemir, Kiymet, et al.
Published: (2025)
FairDeDup: Detecting and Mitigating Vision-Language Fairness Disparities in Semantic Dataset Deduplication
by: Slyman, Eric, et al.
Published: (2024)
by: Slyman, Eric, et al.
Published: (2024)
The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers
by: Qi, Daiqing, et al.
Published: (2025)
by: Qi, Daiqing, et al.
Published: (2025)
Are Bias Mitigation Techniques for Deep Learning Effective?
by: Shrestha, Robik, et al.
Published: (2021)
by: Shrestha, Robik, et al.
Published: (2021)
OccamNets: Mitigating Dataset Bias by Favoring Simpler Hypotheses
by: Shrestha, Robik, et al.
Published: (2022)
by: Shrestha, Robik, et al.
Published: (2022)
SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner
by: Zhou, Yufan, et al.
Published: (2024)
by: Zhou, Yufan, et al.
Published: (2024)
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
by: Zhang, Yanzhe, et al.
Published: (2023)
by: Zhang, Yanzhe, et al.
Published: (2023)
Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents
by: Al-Tawaha, Ahmad, et al.
Published: (2026)
by: Al-Tawaha, Ahmad, et al.
Published: (2026)
Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMs
by: Kang, Feiyang, et al.
Published: (2024)
by: Kang, Feiyang, et al.
Published: (2024)
SOHES: Self-supervised Open-world Hierarchical Entity Segmentation
by: Cao, Shengcao, et al.
Published: (2024)
by: Cao, Shengcao, et al.
Published: (2024)
ARTIST: Improving the Generation of Text-rich Images with Disentangled Diffusion Models and Large Language Models
by: Zhang, Jianyi, et al.
Published: (2024)
by: Zhang, Jianyi, et al.
Published: (2024)
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
by: Slyman, Eric, et al.
Published: (2025)
by: Slyman, Eric, et al.
Published: (2025)
Vision Transformers Need More Than Registers
by: Shi, Cheng, et al.
Published: (2026)
by: Shi, Cheng, et al.
Published: (2026)
MMR: Evaluating Reading Ability of Large Multimodal Models
by: Chen, Jian, et al.
Published: (2024)
by: Chen, Jian, et al.
Published: (2024)
LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models
by: Zhang, Ruiyi, et al.
Published: (2024)
by: Zhang, Ruiyi, et al.
Published: (2024)
Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder
by: Li, Siting, et al.
Published: (2024)
by: Li, Siting, et al.
Published: (2024)
More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models
by: Song, Xurui, et al.
Published: (2025)
by: Song, Xurui, et al.
Published: (2025)
Towards Visual Text Grounding of Multimodal Large Language Model
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Aggregate-Combine-Readout GNNs Are More Expressive Than Logic C2
by: Hauke, Stan P, et al.
Published: (2025)
by: Hauke, Stan P, et al.
Published: (2025)
Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?
by: Chung, Jiwan, et al.
Published: (2025)
by: Chung, Jiwan, et al.
Published: (2025)
TextLap: Customizing Language Models for Text-to-Layout Planning
by: Chen, Jian, et al.
Published: (2024)
by: Chen, Jian, et al.
Published: (2024)
TRINS: Towards Multimodal Language Models that Can Read
by: Zhang, Ruiyi, et al.
Published: (2024)
by: Zhang, Ruiyi, et al.
Published: (2024)
A Solver-in-the-Loop Framework for Improving LLMs on Answer Set Programming for Logic Puzzle Solving
by: Schrader, Timo Pierre, et al.
Published: (2025)
by: Schrader, Timo Pierre, et al.
Published: (2025)
They're All Doctors: Synthesizing Diverse Counterfactuals to Mitigate Associative Bias
by: Magid, Salma Abdel, et al.
Published: (2024)
by: Magid, Salma Abdel, et al.
Published: (2024)
Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency
by: Liu, Junming, et al.
Published: (2026)
by: Liu, Junming, et al.
Published: (2026)
VipAct: Visual-Perception Enhancement via Specialized VLM Agent Collaboration and Tool-use
by: Zhang, Zhehao, et al.
Published: (2024)
by: Zhang, Zhehao, et al.
Published: (2024)
Similar Items
-
Seeing Through Words: Controlling Visual Retrieval Quality with Language Models
by: Lu, Jianglin, et al.
Published: (2026) -
Improving Large Vision and Language Models by Learning from a Panel of Peers
by: Hernandez, Jefferson, et al.
Published: (2025) -
MAGNET: Augmenting Generative Decoders with Representation Learning and Infilling Capabilities
by: Khosla, Savya, et al.
Published: (2025) -
RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward
by: Wu, Qiucheng, et al.
Published: (2026) -
Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
by: Yang, Ziyan, et al.
Published: (2022)