Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Ge, Yuyao, Liu, Shenghua, Wang, Yiwei, Mei, Lingrui, Bi, Baolong, Zhou, Xuanshan, Yao, Jiayu, Guo, Jiafeng, Cheng, Xueqi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HighlightBench: Benchmarking Markup-Driven Table Reasoning in Scientific Documents
by: Wang, Lexin, et al.
Published: (2026)
by: Wang, Lexin, et al.
Published: (2026)
Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking
by: Ge, Yuyao, et al.
Published: (2025)
by: Ge, Yuyao, et al.
Published: (2025)
Prism-$Δ$: Differential Subspace Steering for Prompt Highlighting in Large Language Models
by: Ge, Yuyao, et al.
Published: (2026)
by: Ge, Yuyao, et al.
Published: (2026)
Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain Reasoning
by: Bi, Baolong, et al.
Published: (2025)
by: Bi, Baolong, et al.
Published: (2025)
Decoding by Contrasting Knowledge: Enhancing LLMs' Confidence on Edited Facts
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
Who is in the Spotlight: The Hidden Bias Undermining Multimodal Retrieval-Augmented Generation
by: Yao, Jiayu, et al.
Published: (2025)
by: Yao, Jiayu, et al.
Published: (2025)
Gated Differentiable Working Memory for Long-Context Language Modeling
by: Mei, Lingrui, et al.
Published: (2026)
by: Mei, Lingrui, et al.
Published: (2026)
SLANG: New Concept Comprehension of Large Language Models
by: Mei, Lingrui, et al.
Published: (2024)
by: Mei, Lingrui, et al.
Published: (2024)
LPNL: Scalable Link Prediction with Large Language Models
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
Can Graph Descriptive Order Affect Solving Graph Problems with LLMs?
by: Ge, Yuyao, et al.
Published: (2024)
by: Ge, Yuyao, et al.
Published: (2024)
a1: Steep Test-time Scaling Law via Environment Augmented Generation
by: Mei, Lingrui, et al.
Published: (2025)
by: Mei, Lingrui, et al.
Published: (2025)
Not in Sync: Unveiling Temporal Bias in Audio Chat Models
by: Yao, Jiayu, et al.
Published: (2025)
by: Yao, Jiayu, et al.
Published: (2025)
"Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' Jailbreak
by: Mei, Lingrui, et al.
Published: (2024)
by: Mei, Lingrui, et al.
Published: (2024)
HiddenGuard: Fine-Grained Safe Generation with Specialized Representation Router
by: Mei, Lingrui, et al.
Published: (2024)
by: Mei, Lingrui, et al.
Published: (2024)
MVAM: Multi-View Attention Method for Fine-grained Image-Text Matching
by: Cui, Wanqing, et al.
Published: (2024)
by: Cui, Wanqing, et al.
Published: (2024)
Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language Models
by: Bi, Baolong, et al.
Published: (2025)
by: Bi, Baolong, et al.
Published: (2025)
Adaptive Token Biaser: Knowledge Editing via Biasing Key Entities
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
StruEdit: Structured Outputs Enable the Fast and Accurate Knowledge Editing for Large Language Models
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
A Survey of Vibe Coding with Large Language Models
by: Ge, Yuyao, et al.
Published: (2025)
by: Ge, Yuyao, et al.
Published: (2025)
PromptCD: Test-Time Behavior Enhancement via Polarity-Prompt Contrastive Decoding
by: Bi, Baolong, et al.
Published: (2026)
by: Bi, Baolong, et al.
Published: (2026)
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs
by: Bi, Baolong, et al.
Published: (2025)
by: Bi, Baolong, et al.
Published: (2025)
CLIPure: Purification in Latent Space via CLIP for Adversarially Robust Zero-Shot Classification
by: Zhang, Mingkun, et al.
Published: (2025)
by: Zhang, Mingkun, et al.
Published: (2025)
Visual Transformation Telling
by: Cui, Wanqing, et al.
Published: (2023)
by: Cui, Wanqing, et al.
Published: (2023)
TIR-Flow: Active Video Search and Reasoning with Frozen VLMs
by: Jin, Hongbo, et al.
Published: (2026)
by: Jin, Hongbo, et al.
Published: (2026)
Classifier Guidance Enhances Diffusion-based Adversarial Purification by Preserving Predictive Information
by: Zhang, Mingkun, et al.
Published: (2024)
by: Zhang, Mingkun, et al.
Published: (2024)
CausalDiff: Causality-Inspired Disentanglement via Diffusion Model for Adversarial Defense
by: Zhang, Mingkun, et al.
Published: (2024)
by: Zhang, Mingkun, et al.
Published: (2024)
Is Factuality Enhancement a Free Lunch For LLMs? Better Factuality Can Lead to Worse Context-Faithfulness
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
by: Zhang, Xintong, et al.
Published: (2025)
by: Zhang, Xintong, et al.
Published: (2025)
VisualActBench: Can VLMs See and Act like a Human?
by: Zhang, Daoan, et al.
Published: (2025)
by: Zhang, Daoan, et al.
Published: (2025)
Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
by: Pan, Zhiyu, et al.
Published: (2026)
by: Pan, Zhiyu, et al.
Published: (2026)
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
by: Zhang, Qizhe, et al.
Published: (2024)
by: Zhang, Qizhe, et al.
Published: (2024)
Vision Verification Enhanced Fusion of VLMs for Efficient Visual Reasoning
by: Tekin, Selim Furkan, et al.
Published: (2026)
by: Tekin, Selim Furkan, et al.
Published: (2026)
A Survey of Context Engineering for Large Language Models
by: Mei, Lingrui, et al.
Published: (2025)
by: Mei, Lingrui, et al.
Published: (2025)
Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs
by: Li, Yiwei, et al.
Published: (2026)
by: Li, Yiwei, et al.
Published: (2026)
ViFP: A Framework for Visual False Positive Detection to Enhance Reasoning Reliability in VLMs
by: Zhang, Ben, et al.
Published: (2025)
by: Zhang, Ben, et al.
Published: (2025)
Beyond Shortcuts: Mitigating Visual Illusions in Frozen VLMs via Qualitative Reasoning
by: Guo, Hao, et al.
Published: (2026)
by: Guo, Hao, et al.
Published: (2026)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
by: Berman, Shmuel, et al.
Published: (2025)
by: Berman, Shmuel, et al.
Published: (2025)
Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception
by: Ni, Shiyu, et al.
Published: (2025)
by: Ni, Shiyu, et al.
Published: (2025)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
by: Li, Sifan, et al.
Published: (2025)
by: Li, Sifan, et al.
Published: (2025)
Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs
by: Elmansoury, Sary, et al.
Published: (2025)
by: Elmansoury, Sary, et al.
Published: (2025)
Similar Items
-
HighlightBench: Benchmarking Markup-Driven Table Reasoning in Scientific Documents
by: Wang, Lexin, et al.
Published: (2026) -
Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking
by: Ge, Yuyao, et al.
Published: (2025) -
Prism-$Δ$: Differential Subspace Steering for Prompt Highlighting in Large Language Models
by: Ge, Yuyao, et al.
Published: (2026) -
Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain Reasoning
by: Bi, Baolong, et al.
Published: (2025) -
Decoding by Contrasting Knowledge: Enhancing LLMs' Confidence on Edited Facts
by: Bi, Baolong, et al.
Published: (2024)