Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Yingjie, Bai, Xuefeng, Chen, Kehai, Xiang, Yang, Pan, Youcheng, Hou, Yongshuai, Guan, Weili, Yu, Jun, Zhang, Min |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
Evaluating and Steering Modality Preferences in Multimodal Large Language Model
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
Mitigating Translationese Bias in Multilingual LLM-as-a-Judge via Disentangled Information Bottleneck
von: Zhang, Hongbin, et al.
Veröffentlicht: (2026)
von: Zhang, Hongbin, et al.
Veröffentlicht: (2026)
Paying More Attention to Source Context: Mitigating Unfaithful Translations from Large Language Model
von: Zhang, Hongbin, et al.
Veröffentlicht: (2024)
von: Zhang, Hongbin, et al.
Veröffentlicht: (2024)
Mitigating Multimodal Hallucination via Phase-wise Self-reward
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
Exploring the Translation Mechanism of Large Language Models
von: Zhang, Hongbin, et al.
Veröffentlicht: (2025)
von: Zhang, Hongbin, et al.
Veröffentlicht: (2025)
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey
von: Qu, Bingzheng, et al.
Veröffentlicht: (2026)
von: Qu, Bingzheng, et al.
Veröffentlicht: (2026)
Beyond Rigid: Benchmarking Non-Rigid Video Editing
von: Qu, Bingzheng, et al.
Veröffentlicht: (2026)
von: Qu, Bingzheng, et al.
Veröffentlicht: (2026)
Agentic Tool Use in Large Language Models
von: Hu, Jinchao, et al.
Veröffentlicht: (2026)
von: Hu, Jinchao, et al.
Veröffentlicht: (2026)
Beyond Token-Level Policy Gradients for Complex Reasoning with Large Language Models
von: Xu, Mufan, et al.
Veröffentlicht: (2026)
von: Xu, Mufan, et al.
Veröffentlicht: (2026)
SAT: Balancing Reasoning Accuracy and Efficiency with Stepwise Adaptive Thinking
von: Huang, Weiyang, et al.
Veröffentlicht: (2026)
von: Huang, Weiyang, et al.
Veröffentlicht: (2026)
Handling Imbalanced Pseudolabels for Vision-Language Models with Concept Alignment and Confusion-Aware Calibrated Margin
von: Wang, Yuchen, et al.
Veröffentlicht: (2025)
von: Wang, Yuchen, et al.
Veröffentlicht: (2025)
LinguaLIFT: An Effective Two-stage Instruction Tuning Framework for Low-Resource Language Reasoning
von: Zhang, Hongbin, et al.
Veröffentlicht: (2024)
von: Zhang, Hongbin, et al.
Veröffentlicht: (2024)
TF-Attack: Transferable and Fast Adversarial Attacks on Large Language Models
von: Li, Zelin, et al.
Veröffentlicht: (2024)
von: Li, Zelin, et al.
Veröffentlicht: (2024)
Identifying and Mitigating Position Bias of Multi-image Vision-Language Models
von: Tian, Xinyu, et al.
Veröffentlicht: (2025)
von: Tian, Xinyu, et al.
Veröffentlicht: (2025)
Generator-Assistant Stepwise Rollback Framework for Large Language Model Agent
von: Li, Xingzuo, et al.
Veröffentlicht: (2025)
von: Li, Xingzuo, et al.
Veröffentlicht: (2025)
Through the Lens of Character: Resolving Modality-Role Interference in Multimodal Role-Playing Agent
von: Tang, Yihong, et al.
Veröffentlicht: (2026)
von: Tang, Yihong, et al.
Veröffentlicht: (2026)
The Power of Personality: A Human Simulation Perspective to Investigate Large Language Model Agents
von: Duan, Yifan, et al.
Veröffentlicht: (2025)
von: Duan, Yifan, et al.
Veröffentlicht: (2025)
VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
von: Wang, Zihu, et al.
Veröffentlicht: (2025)
von: Wang, Zihu, et al.
Veröffentlicht: (2025)
Beyond Unimodal Shortcuts: MLLMs as Cross-Modal Reasoners for Grounded Named Entity Recognition
von: Ma, Jinlong, et al.
Veröffentlicht: (2026)
von: Ma, Jinlong, et al.
Veröffentlicht: (2026)
Efficient Safety Alignment of Large Language Models via Preference Re-ranking and Representation-based Reward Modeling
von: Deng, Qiyuan, et al.
Veröffentlicht: (2025)
von: Deng, Qiyuan, et al.
Veröffentlicht: (2025)
Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment
von: Zhang, Hongbin, et al.
Veröffentlicht: (2025)
von: Zhang, Hongbin, et al.
Veröffentlicht: (2025)
Towards Mitigating Modality Bias in Vision-Language Models for Temporal Action Localization
von: Li, Jiaqi, et al.
Veröffentlicht: (2026)
von: Li, Jiaqi, et al.
Veröffentlicht: (2026)
Large Language Models for Classical Chinese Poetry Translation: Benchmarking, Evaluating, and Improving
von: Chen, Andong, et al.
Veröffentlicht: (2024)
von: Chen, Andong, et al.
Veröffentlicht: (2024)
Blind to Position, Biased in Language: Probing Mid-Layer Representational Bias in Vision-Language Encoders for Zero-Shot Language-Grounded Spatial Understanding
von: An, Na Min, et al.
Veröffentlicht: (2025)
von: An, Na Min, et al.
Veröffentlicht: (2025)
DUAL-REFLECT: Enhancing Large Language Models for Reflective Translation through Dual Learning Feedback Mechanisms
von: Chen, Andong, et al.
Veröffentlicht: (2024)
von: Chen, Andong, et al.
Veröffentlicht: (2024)
Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language Models
von: Atabuzzaman, Md., et al.
Veröffentlicht: (2025)
von: Atabuzzaman, Md., et al.
Veröffentlicht: (2025)
Investigating Spatial Attention Bias in Vision-Language Models
von: Chaudhary, Aryan, et al.
Veröffentlicht: (2025)
von: Chaudhary, Aryan, et al.
Veröffentlicht: (2025)
debiaSAE: Benchmarking and Mitigating Vision-Language Model Bias
von: Sasse, Kuleen, et al.
Veröffentlicht: (2024)
von: Sasse, Kuleen, et al.
Veröffentlicht: (2024)
XIFBench: Evaluating Large Language Models on Multilingual Instruction Following
von: Li, Zhenyu, et al.
Veröffentlicht: (2025)
von: Li, Zhenyu, et al.
Veröffentlicht: (2025)
Instruction Anchor: Dissecting the Mechanistic Dynamics of Modality Arbitration
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
HITSZ's End-To-End Speech Translation Systems Combining Sequence-to-Sequence Auto Speech Recognition Model and Indic Large Language Model for IWSLT 2025 in Indic Track
von: Wei, Xuchen, et al.
Veröffentlicht: (2025)
von: Wei, Xuchen, et al.
Veröffentlicht: (2025)
Mitigating Multilingual Hallucination in Large Vision-Language Models
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
YARD: Y-Architecture Register Decoding for Efficient Hallucination Mitigation in Large Vision-Language Models
von: Chen, Ting, et al.
Veröffentlicht: (2026)
von: Chen, Ting, et al.
Veröffentlicht: (2026)
Effective and Efficient Adversarial Detection for Vision-Language Models via A Single Vector
von: Huang, Youcheng, et al.
Veröffentlicht: (2024)
von: Huang, Youcheng, et al.
Veröffentlicht: (2024)
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
von: Wang, Xintong, et al.
Veröffentlicht: (2024)
von: Wang, Xintong, et al.
Veröffentlicht: (2024)
Question-guided Knowledge Graph Re-scoring and Injection for Knowledge Graph Question Answering
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
A Survey on Human Preference Learning for Large Language Models
von: Jiang, Ruili, et al.
Veröffentlicht: (2024)
von: Jiang, Ruili, et al.
Veröffentlicht: (2024)
Beyond Semantics: Rediscovering Spatial Awareness in Vision-Language Models
von: Qi, Jianing, et al.
Veröffentlicht: (2025)
von: Qi, Jianing, et al.
Veröffentlicht: (2025)
Hierarchical Pre-Training of Vision Encoders with Large Language Models
von: Lee, Eugene, et al.
Veröffentlicht: (2026)
von: Lee, Eugene, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024) -
Evaluating and Steering Modality Preferences in Multimodal Large Language Model
von: Zhang, Yu, et al.
Veröffentlicht: (2025) -
Mitigating Translationese Bias in Multilingual LLM-as-a-Judge via Disentangled Information Bottleneck
von: Zhang, Hongbin, et al.
Veröffentlicht: (2026) -
Paying More Attention to Source Context: Mitigating Unfaithful Translations from Large Language Model
von: Zhang, Hongbin, et al.
Veröffentlicht: (2024) -
Mitigating Multimodal Hallucination via Phase-wise Self-reward
von: Zhang, Yu, et al.
Veröffentlicht: (2026)