A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Balasubramanian, Sriram, Basu, Samyadeep, Feizi, Soheil |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model
por: Zhang, Sinin, et al.
Publicado: (2026)
por: Zhang, Sinin, et al.
Publicado: (2026)
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
por: Le, Van-Truong
Publicado: (2026)
por: Le, Van-Truong
Publicado: (2026)
On Measuring Faithfulness or Self-consistency of Natural Language Explanations
por: Parcalabescu, Letitia, et al.
Publicado: (2023)
por: Parcalabescu, Letitia, et al.
Publicado: (2023)
Universal Adversarial Attack on Aligned Multimodal LLMs
por: Rahmatullaev, Temurbek, et al.
Publicado: (2025)
por: Rahmatullaev, Temurbek, et al.
Publicado: (2025)
Large Language Model for Qualitative Research -- A Systematic Mapping Study
por: Barros, Cauã Ferreira, et al.
Publicado: (2024)
por: Barros, Cauã Ferreira, et al.
Publicado: (2024)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
On the Limitations of Vision-Language Models in Understanding Image Transforms
por: Anis, Ahmad Mustafa, et al.
Publicado: (2025)
por: Anis, Ahmad Mustafa, et al.
Publicado: (2025)
Evaluating Voice Command Pipelines for Drone Control: From STT and LLM to Direct Classification and Siamese Networks
por: Simões, Lucca Emmanuel Pineli, et al.
Publicado: (2024)
por: Simões, Lucca Emmanuel Pineli, et al.
Publicado: (2024)
CORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships?
por: Ramakrishnan, Aashish Anantha, et al.
Publicado: (2025)
por: Ramakrishnan, Aashish Anantha, et al.
Publicado: (2025)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
por: Kim, Soyeon, et al.
Publicado: (2026)
por: Kim, Soyeon, et al.
Publicado: (2026)
Bounding Hallucinations: Information-Theoretic Guarantees for RAG Systems via Merlin-Arthur Protocols
por: Deiseroth, Björn, et al.
Publicado: (2025)
por: Deiseroth, Björn, et al.
Publicado: (2025)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
por: Dai, Song, et al.
Publicado: (2025)
por: Dai, Song, et al.
Publicado: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
por: Tong, Jingqi, et al.
Publicado: (2025)
por: Tong, Jingqi, et al.
Publicado: (2025)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
por: Bian, Zhipeng, et al.
Publicado: (2026)
por: Bian, Zhipeng, et al.
Publicado: (2026)
Beyond Visual Understanding: Introducing PARROT-360V for Vision Language Model Benchmarking
por: Khurdula, Harsha Vardhan, et al.
Publicado: (2024)
por: Khurdula, Harsha Vardhan, et al.
Publicado: (2024)
Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
por: Ji, Binbin, et al.
Publicado: (2025)
por: Ji, Binbin, et al.
Publicado: (2025)
Step-CoT: Stepwise Visual Chain-of-Thought for Medical Visual Question Answering
por: Fan, Lin, et al.
Publicado: (2026)
por: Fan, Lin, et al.
Publicado: (2026)
Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs
por: Asanuma, Haruka, et al.
Publicado: (2025)
por: Asanuma, Haruka, et al.
Publicado: (2025)
MIRA: Empowering One-Touch AI Services on Smartphones with MLLM-based Instruction Recommendation
por: Bian, Zhipeng, et al.
Publicado: (2025)
por: Bian, Zhipeng, et al.
Publicado: (2025)
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
ReaderLM-v2: Small Language Model for HTML to Markdown and JSON
por: Wang, Feng, et al.
Publicado: (2025)
por: Wang, Feng, et al.
Publicado: (2025)
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
por: Rudman, William, et al.
Publicado: (2026)
por: Rudman, William, et al.
Publicado: (2026)
Analyzing Quality, Bias, and Performance in Text-to-Image Generative Models
por: Masrourisaadat, Nila, et al.
Publicado: (2024)
por: Masrourisaadat, Nila, et al.
Publicado: (2024)
Growing Perspectives: Modelling Embodied Perspective Taking and Inner Narrative Development Using Large Language Models
por: Patania, Sabrina, et al.
Publicado: (2025)
por: Patania, Sabrina, et al.
Publicado: (2025)
Context-Dependent Affordance Computation in Vision-Language Models
por: Farzulla, Murad
Publicado: (2026)
por: Farzulla, Murad
Publicado: (2026)
MemeCraft: Contextual and Stance-Driven Multimodal Meme Generation
por: Wang, Han, et al.
Publicado: (2024)
por: Wang, Han, et al.
Publicado: (2024)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
Who Sees What? Structured Thought-Action Sequences for Epistemic Reasoning in LLMs
por: Annese, Luca, et al.
Publicado: (2025)
por: Annese, Luca, et al.
Publicado: (2025)
VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models
por: Cao, Jingtao, et al.
Publicado: (2024)
por: Cao, Jingtao, et al.
Publicado: (2024)
Leum-VL Technical Report
por: He, Yuxuan, et al.
Publicado: (2026)
por: He, Yuxuan, et al.
Publicado: (2026)
Learning the meanings of function words from grounded language using a visual question answering model
por: Portelance, Eva, et al.
Publicado: (2023)
por: Portelance, Eva, et al.
Publicado: (2023)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
por: Yang, Shan
Publicado: (2026)
por: Yang, Shan
Publicado: (2026)
Relative Drawing Identification Complexity is Invariant to Modality in Vision-Language Models
por: Freitas, Diogo, et al.
Publicado: (2025)
por: Freitas, Diogo, et al.
Publicado: (2025)
More Than Meets the Eye: Measuring the Semiotic Gap in Vision-Language Models via Semantic Anchorage
por: He, Wei
Publicado: (2026)
por: He, Wei
Publicado: (2026)
Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
por: Ji, Yikun, et al.
Publicado: (2025)
por: Ji, Yikun, et al.
Publicado: (2025)
InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
por: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Publicado: (2025)
por: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Publicado: (2025)
TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning
por: Sanders, Kate, et al.
Publicado: (2024)
por: Sanders, Kate, et al.
Publicado: (2024)
IRONIC: Coherence-Aware Reasoning Chains for Multi-Modal Sarcasm Detection
por: Ramakrishnan, Aashish Anantha, et al.
Publicado: (2025)
por: Ramakrishnan, Aashish Anantha, et al.
Publicado: (2025)
DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
por: Hou, Zhiyi, et al.
Publicado: (2025)
por: Hou, Zhiyi, et al.
Publicado: (2025)
Defending against Backdoor Attacks via Module Switching
por: Li, Weijun, et al.
Publicado: (2025)
por: Li, Weijun, et al.
Publicado: (2025)
Ejemplares similares
-
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model
por: Zhang, Sinin, et al.
Publicado: (2026) -
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
por: Le, Van-Truong
Publicado: (2026) -
On Measuring Faithfulness or Self-consistency of Natural Language Explanations
por: Parcalabescu, Letitia, et al.
Publicado: (2023) -
Universal Adversarial Attack on Aligned Multimodal LLMs
por: Rahmatullaev, Temurbek, et al.
Publicado: (2025) -
Large Language Model for Qualitative Research -- A Systematic Mapping Study
por: Barros, Cauã Ferreira, et al.
Publicado: (2024)