Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Xiaoyuan, Wang, Wenxuan, Yuan, Youliang, Huang, Jen-tse, Liu, Qiuzhi, He, Pinjia, Tu, Zhaopeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
por: Wang, Wenxuan, et al.
Publicado: (2025)
por: Wang, Wenxuan, et al.
Publicado: (2025)
GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
por: Yuan, Youliang, et al.
Publicado: (2023)
por: Yuan, Youliang, et al.
Publicado: (2023)
Chain-of-Jailbreak Attack for Image Generation Models via Editing Step by Step
por: Wang, Wenxuan, et al.
Publicado: (2024)
por: Wang, Wenxuan, et al.
Publicado: (2024)
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
por: Huang, Jen-Tse, et al.
Publicado: (2025)
por: Huang, Jen-Tse, et al.
Publicado: (2025)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
por: Yuan, Youliang, et al.
Publicado: (2024)
por: Yuan, Youliang, et al.
Publicado: (2024)
VisBias: Measuring Explicit and Implicit Social Biases in Vision Language Models
por: Huang, Jen-tse, et al.
Publicado: (2025)
por: Huang, Jen-tse, et al.
Publicado: (2025)
Towards Evaluating Proactive Risk Awareness of Multimodal Language Models
por: Yuan, Youliang, et al.
Publicado: (2025)
por: Yuan, Youliang, et al.
Publicado: (2025)
AI Sees Your Location, But With A Bias Toward The Wealthy World
por: Huang, Jingyuan, et al.
Publicado: (2025)
por: Huang, Jingyuan, et al.
Publicado: (2025)
Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
por: Yuan, Youliang, et al.
Publicado: (2025)
por: Yuan, Youliang, et al.
Publicado: (2025)
All Languages Matter: On the Multilingual Safety of Large Language Models
por: Wang, Wenxuan, et al.
Publicado: (2023)
por: Wang, Wenxuan, et al.
Publicado: (2023)
Who is ChatGPT? Benchmarking LLMs' Psychological Portrayal Using PsychoBench
por: Huang, Jen-tse, et al.
Publicado: (2023)
por: Huang, Jen-tse, et al.
Publicado: (2023)
Difficult Task Yes but Simple Task No: Unveiling the Laziness in Multimodal LLMs
por: Zhao, Sihang, et al.
Publicado: (2024)
por: Zhao, Sihang, et al.
Publicado: (2024)
LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models
por: Wan, Yuxuan, et al.
Publicado: (2024)
por: Wan, Yuxuan, et al.
Publicado: (2024)
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
por: Huang, Jen-tse, et al.
Publicado: (2024)
por: Huang, Jen-tse, et al.
Publicado: (2024)
On the Shortcut Learning in Multilingual Neural Machine Translation
por: Wang, Wenxuan, et al.
Publicado: (2024)
por: Wang, Wenxuan, et al.
Publicado: (2024)
New Job, New Gender? Measuring the Social Bias in Image Generation Models
por: Wang, Wenxuan, et al.
Publicado: (2024)
por: Wang, Wenxuan, et al.
Publicado: (2024)
Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models
por: Zhu, Tinghui, et al.
Publicado: (2024)
por: Zhu, Tinghui, et al.
Publicado: (2024)
SHAPE: Unifying Safety, Helpfulness and Pedagogy for Educational LLMs
por: Zhao, Sihang, et al.
Publicado: (2026)
por: Zhao, Sihang, et al.
Publicado: (2026)
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
por: Liu, Xiaoyuan, et al.
Publicado: (2025)
por: Liu, Xiaoyuan, et al.
Publicado: (2025)
Emotionally Numb or Empathetic? Evaluating How LLMs Feel Using EmotionBench
por: Huang, Jen-tse, et al.
Publicado: (2023)
por: Huang, Jen-tse, et al.
Publicado: (2023)
Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language Models
por: Wang, Wenxuan, et al.
Publicado: (2023)
por: Wang, Wenxuan, et al.
Publicado: (2023)
PaSBench-Video: A Streaming Video Benchmark for Proactive Safety Warning
por: Zhao, Yusong, et al.
Publicado: (2026)
por: Zhao, Yusong, et al.
Publicado: (2026)
Identifying the Achilles' Heel: An Iterative Method for Dynamically Uncovering Factual Errors in Large Language Models
por: Wang, Wenxuan, et al.
Publicado: (2024)
por: Wang, Wenxuan, et al.
Publicado: (2024)
Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
por: Wang, Yue, et al.
Publicado: (2025)
por: Wang, Yue, et al.
Publicado: (2025)
Vision Enhancing LLMs: Empowering Multimodal Knowledge Storage and Sharing in LLMs
por: Li, Yunxin, et al.
Publicado: (2023)
por: Li, Yunxin, et al.
Publicado: (2023)
Probing Multimodal Large Language Models on Cognitive Biases in Chinese Short-Video Misinformation
por: Huang, Jen-tse, et al.
Publicado: (2026)
por: Huang, Jen-tse, et al.
Publicado: (2026)
On the Failure of Latent State Persistence in Large Language Models
por: Huang, Jen-tse, et al.
Publicado: (2025)
por: Huang, Jen-tse, et al.
Publicado: (2025)
From Sight to Insight: Improving Visual Reasoning Capabilities of Multimodal Models via Reinforcement Learning
por: Sharif, Omar, et al.
Publicado: (2026)
por: Sharif, Omar, et al.
Publicado: (2026)
The Hunger Game Debate: On the Emergence of Over-Competition in Multi-Agent Systems
por: Ma, Xinbei, et al.
Publicado: (2025)
por: Ma, Xinbei, et al.
Publicado: (2025)
How Well Can LLMs Echo Us? Evaluating AI Chatbots' Role-Play Ability with ECHO
por: Ng, Man Tik, et al.
Publicado: (2024)
por: Ng, Man Tik, et al.
Publicado: (2024)
The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models
por: Ji, Ke, et al.
Publicado: (2025)
por: Ji, Ke, et al.
Publicado: (2025)
Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
por: Wang, Mengru, et al.
Publicado: (2025)
por: Wang, Mengru, et al.
Publicado: (2025)
FairCoder: Evaluating Social Bias of LLMs in Code Generation
por: Du, Yongkang, et al.
Publicado: (2025)
por: Du, Yongkang, et al.
Publicado: (2025)
ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention
por: Liu, Wenjie, et al.
Publicado: (2026)
por: Liu, Wenjie, et al.
Publicado: (2026)
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
por: Jiang, Houcheng, et al.
Publicado: (2026)
por: Jiang, Houcheng, et al.
Publicado: (2026)
ComboBench: Can LLMs Manipulate Physical Devices to Play Virtual Reality Games?
por: Li, Shuqing, et al.
Publicado: (2025)
por: Li, Shuqing, et al.
Publicado: (2025)
Can LLMs Grasp Implicit Cultural Values? Benchmarking LLMs' Cultural Intelligence with CQ-Bench
por: Liu, Ziyi, et al.
Publicado: (2025)
por: Liu, Ziyi, et al.
Publicado: (2025)
Conflict Adaptation in Vision-Language Models
por: Hu, Xiaoyang
Publicado: (2025)
por: Hu, Xiaoyang
Publicado: (2025)
A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?
por: Chen, Ada, et al.
Publicado: (2025)
por: Chen, Ada, et al.
Publicado: (2025)
DeepSight: Bridging Depth Maps and Language with a Depth-Driven Multimodal Model
por: Yang, Hao, et al.
Publicado: (2026)
por: Yang, Hao, et al.
Publicado: (2026)
Ejemplares similares
-
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
por: Wang, Wenxuan, et al.
Publicado: (2025) -
GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
por: Yuan, Youliang, et al.
Publicado: (2023) -
Chain-of-Jailbreak Attack for Image Generation Models via Editing Step by Step
por: Wang, Wenxuan, et al.
Publicado: (2024) -
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
por: Huang, Jen-Tse, et al.
Publicado: (2025) -
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
por: Yuan, Youliang, et al.
Publicado: (2024)