Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Kamath, Amita, Hessel, Jack, Chandu, Khyathi, Hwang, Jena D., Chang, Kai-Wei, Krishna, Ranjay |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Hard Positive Truth about Vision-Language Compositionality
di: Kamath, Amita, et al.
Pubblicazione: (2024)
di: Kamath, Amita, et al.
Pubblicazione: (2024)
Selective "Selective Prediction": Reducing Unnecessary Abstention in Vision-Language Reasoning
di: Srinivasan, Tejas, et al.
Pubblicazione: (2024)
di: Srinivasan, Tejas, et al.
Pubblicazione: (2024)
Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models
di: Sicilia, Anthony, et al.
Pubblicazione: (2024)
di: Sicilia, Anthony, et al.
Pubblicazione: (2024)
Visual Representations inside the Language Model
di: Liu, Benlin, et al.
Pubblicazione: (2025)
di: Liu, Benlin, et al.
Pubblicazione: (2025)
GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
di: Kamath, Amita, et al.
Pubblicazione: (2025)
di: Kamath, Amita, et al.
Pubblicazione: (2025)
Semantic and Expressive Variation in Image Captions Across Languages
di: Ye, Andre, et al.
Pubblicazione: (2023)
di: Ye, Andre, et al.
Pubblicazione: (2023)
Matryoshka Query Transformer for Large Vision-Language Models
di: Hu, Wenbo, et al.
Pubblicazione: (2024)
di: Hu, Wenbo, et al.
Pubblicazione: (2024)
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
di: Chandu, Khyathi Raghavi, et al.
Pubblicazione: (2024)
di: Chandu, Khyathi Raghavi, et al.
Pubblicazione: (2024)
Agent Lumos: Unified and Modular Training for Open-Source Language Agents
di: Yin, Da, et al.
Pubblicazione: (2023)
di: Yin, Da, et al.
Pubblicazione: (2023)
Vision-Language Models Can't See the Obvious
di: Dahou, Yasser, et al.
Pubblicazione: (2025)
di: Dahou, Yasser, et al.
Pubblicazione: (2025)
L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects
di: Yamada, Yutaro, et al.
Pubblicazione: (2024)
di: Yamada, Yutaro, et al.
Pubblicazione: (2024)
"Flex Tape Can't Fix That": Bias and Misinformation in Edited Language Models
di: Halevy, Karina, et al.
Pubblicazione: (2024)
di: Halevy, Karina, et al.
Pubblicazione: (2024)
Exposing and Addressing Cross-Task Inconsistency in Unified Vision-Language Models
di: Maharana, Adyasha, et al.
Pubblicazione: (2023)
di: Maharana, Adyasha, et al.
Pubblicazione: (2023)
UNcommonsense Reasoning: Abductive Reasoning about Uncommon Situations
di: Zhao, Wenting, et al.
Pubblicazione: (2023)
di: Zhao, Wenting, et al.
Pubblicazione: (2023)
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
di: Zhang, Tianyi, et al.
Pubblicazione: (2026)
di: Zhang, Tianyi, et al.
Pubblicazione: (2026)
The Art of Saying No: Contextual Noncompliance in Language Models
di: Brahman, Faeze, et al.
Pubblicazione: (2024)
di: Brahman, Faeze, et al.
Pubblicazione: (2024)
Symbolic Chain-of-Thought Distillation: Small Models Can Also "Think" Step-by-Step
di: Li, Liunian Harold, et al.
Pubblicazione: (2023)
di: Li, Liunian Harold, et al.
Pubblicazione: (2023)
Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models
di: Hu, Yushi, et al.
Pubblicazione: (2023)
di: Hu, Yushi, et al.
Pubblicazione: (2023)
Iterated Learning Improves Compositionality in Large Vision-Language Models
di: Zheng, Chenhao, et al.
Pubblicazione: (2024)
di: Zheng, Chenhao, et al.
Pubblicazione: (2024)
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
di: Lee, Heekyung, et al.
Pubblicazione: (2025)
di: Lee, Heekyung, et al.
Pubblicazione: (2025)
RESTOR: Knowledge Recovery in Machine Unlearning
di: Rezaei, Keivan, et al.
Pubblicazione: (2024)
di: Rezaei, Keivan, et al.
Pubblicazione: (2024)
Can't say cant? Measuring and Reasoning of Dark Jargons in Large Language Models
di: Ji, Xu, et al.
Pubblicazione: (2024)
di: Ji, Xu, et al.
Pubblicazione: (2024)
Semantic Deception: When Reasoning Models Can't Compute an Addition
di: de Leeuw, Nathaniël, et al.
Pubblicazione: (2025)
di: de Leeuw, Nathaniël, et al.
Pubblicazione: (2025)
Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
di: Chung, Jiwan, et al.
Pubblicazione: (2024)
di: Chung, Jiwan, et al.
Pubblicazione: (2024)
Synthetic Visual Genome
di: Park, Jae Sung, et al.
Pubblicazione: (2025)
di: Park, Jae Sung, et al.
Pubblicazione: (2025)
Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty
di: Zhou, Kaitlyn, et al.
Pubblicazione: (2024)
di: Zhou, Kaitlyn, et al.
Pubblicazione: (2024)
Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search
di: Yen, Howard, et al.
Pubblicazione: (2025)
di: Yen, Howard, et al.
Pubblicazione: (2025)
On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training
di: Wu, Xueqing, et al.
Pubblicazione: (2026)
di: Wu, Xueqing, et al.
Pubblicazione: (2026)
AbsenceBench: Language Models Can't Tell What's Missing
di: Fu, Harvey Yiyun, et al.
Pubblicazione: (2025)
di: Fu, Harvey Yiyun, et al.
Pubblicazione: (2025)
VLS: Steering Pretrained Robot Policies via Vision-Language Models
di: Liu, Shuo, et al.
Pubblicazione: (2026)
di: Liu, Shuo, et al.
Pubblicazione: (2026)
The Machine Can't Replace the Human Heart
di: Lin, Baihan
Pubblicazione: (2024)
di: Lin, Baihan
Pubblicazione: (2024)
You Can't Fight in Here! This is BBS!
di: Futrell, Richard, et al.
Pubblicazione: (2026)
di: Futrell, Richard, et al.
Pubblicazione: (2026)
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
di: Deng, Yihe, et al.
Pubblicazione: (2025)
di: Deng, Yihe, et al.
Pubblicazione: (2025)
Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
di: Prakash, Nirmalendu, et al.
Pubblicazione: (2025)
di: Prakash, Nirmalendu, et al.
Pubblicazione: (2025)
BLINK: Multimodal Large Language Models Can See but Not Perceive
di: Fu, Xingyu, et al.
Pubblicazione: (2024)
di: Fu, Xingyu, et al.
Pubblicazione: (2024)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
di: Sun, Yiyou, et al.
Pubblicazione: (2025)
di: Sun, Yiyou, et al.
Pubblicazione: (2025)
IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models
di: You, Haoxuan, et al.
Pubblicazione: (2023)
di: You, Haoxuan, et al.
Pubblicazione: (2023)
Multilingual Diversity Improves Vision-Language Representations
di: Nguyen, Thao, et al.
Pubblicazione: (2024)
di: Nguyen, Thao, et al.
Pubblicazione: (2024)
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
di: Che, Liwei, et al.
Pubblicazione: (2026)
di: Che, Liwei, et al.
Pubblicazione: (2026)
Promptriever: Instruction-Trained Retrievers Can Be Prompted Like Language Models
di: Weller, Orion, et al.
Pubblicazione: (2024)
di: Weller, Orion, et al.
Pubblicazione: (2024)
Documenti analoghi
-
The Hard Positive Truth about Vision-Language Compositionality
di: Kamath, Amita, et al.
Pubblicazione: (2024) -
Selective "Selective Prediction": Reducing Unnecessary Abstention in Vision-Language Reasoning
di: Srinivasan, Tejas, et al.
Pubblicazione: (2024) -
Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models
di: Sicilia, Anthony, et al.
Pubblicazione: (2024) -
Visual Representations inside the Language Model
di: Liu, Benlin, et al.
Pubblicazione: (2025) -
GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
di: Kamath, Amita, et al.
Pubblicazione: (2025)