Selective "Selective Prediction": Reducing Unnecessary Abstention in Vision-Language Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Srinivasan, Tejas, Hessel, Jack, Gupta, Tanmay, Lin, Bill Yuchen, Choi, Yejin, Thomason, Jesse, Chandu, Khyathi Raghavi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
von: Chandu, Khyathi Raghavi, et al.
Veröffentlicht: (2024)
von: Chandu, Khyathi Raghavi, et al.
Veröffentlicht: (2024)
L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects
von: Yamada, Yutaro, et al.
Veröffentlicht: (2024)
von: Yamada, Yutaro, et al.
Veröffentlicht: (2024)
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
von: Kamath, Amita, et al.
Veröffentlicht: (2026)
von: Kamath, Amita, et al.
Veröffentlicht: (2026)
Adjust for Trust: Mitigating Trust-Induced Inappropriate Reliance on AI Assistance
von: Srinivasan, Tejas, et al.
Veröffentlicht: (2025)
von: Srinivasan, Tejas, et al.
Veröffentlicht: (2025)
Agent Lumos: Unified and Modular Training for Open-Source Language Agents
von: Yin, Da, et al.
Veröffentlicht: (2023)
von: Yin, Da, et al.
Veröffentlicht: (2023)
WinoViz: Probing Visual Properties of Objects Under Different States
von: Jin, Woojeong, et al.
Veröffentlicht: (2024)
von: Jin, Woojeong, et al.
Veröffentlicht: (2024)
Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
von: He, Keyu, et al.
Veröffentlicht: (2025)
von: He, Keyu, et al.
Veröffentlicht: (2025)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
RESTOR: Knowledge Recovery in Machine Unlearning
von: Rezaei, Keivan, et al.
Veröffentlicht: (2024)
von: Rezaei, Keivan, et al.
Veröffentlicht: (2024)
From Calibration to Collaboration: LLM Uncertainty Quantification Should Be More Human-Centered
von: Devic, Siddartha, et al.
Veröffentlicht: (2025)
von: Devic, Siddartha, et al.
Veröffentlicht: (2025)
Continual Dialogue State Tracking via Example-Guided Question Answering
von: Cho, Hyundong, et al.
Veröffentlicht: (2023)
von: Cho, Hyundong, et al.
Veröffentlicht: (2023)
PDDL-Mind: Large Language Models are Capable on Belief Reasoning with Reliable State Tracking
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2026)
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2026)
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
The Art of Saying No: Contextual Noncompliance in Language Models
von: Brahman, Faeze, et al.
Veröffentlicht: (2024)
von: Brahman, Faeze, et al.
Veröffentlicht: (2024)
Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
von: Chung, Jiwan, et al.
Veröffentlicht: (2024)
von: Chung, Jiwan, et al.
Veröffentlicht: (2024)
AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
von: Lu, Ximing, et al.
Veröffentlicht: (2024)
von: Lu, Ximing, et al.
Veröffentlicht: (2024)
WildChat: 1M ChatGPT Interaction Logs in the Wild
von: Zhao, Wenting, et al.
Veröffentlicht: (2024)
von: Zhao, Wenting, et al.
Veröffentlicht: (2024)
PSALM-V: Automating Symbolic Planning in Interactive Visual Environments with Large Language Models
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2025)
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2025)
UNcommonsense Reasoning: Abductive Reasoning about Uncommon Situations
von: Zhao, Wenting, et al.
Veröffentlicht: (2023)
von: Zhao, Wenting, et al.
Veröffentlicht: (2023)
MASH: Modeling Abstention via Selective Help-Seeking
von: Gul, Mustafa Omer, et al.
Veröffentlicht: (2025)
von: Gul, Mustafa Omer, et al.
Veröffentlicht: (2025)
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2025)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2025)
Words that make SENSE: Sensorimotor Norms in Learned Lexical Token Representations
von: Gupta, Abhinav, et al.
Veröffentlicht: (2026)
von: Gupta, Abhinav, et al.
Veröffentlicht: (2026)
Symbolic Chain-of-Thought Distillation: Small Models Can Also "Think" Step-by-Step
von: Li, Liunian Harold, et al.
Veröffentlicht: (2023)
von: Li, Liunian Harold, et al.
Veröffentlicht: (2023)
Characterizing LLM Abstention Behavior in Science QA with Context Perturbations
von: Wen, Bingbing, et al.
Veröffentlicht: (2024)
von: Wen, Bingbing, et al.
Veröffentlicht: (2024)
Know Your Limits: A Survey of Abstention in Large Language Models
von: Wen, Bingbing, et al.
Veröffentlicht: (2024)
von: Wen, Bingbing, et al.
Veröffentlicht: (2024)
WildVis: Open Source Visualizer for Million-Scale Chat Logs in the Wild
von: Deng, Yuntian, et al.
Veröffentlicht: (2024)
von: Deng, Yuntian, et al.
Veröffentlicht: (2024)
WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries
von: Zhao, Wenting, et al.
Veröffentlicht: (2024)
von: Zhao, Wenting, et al.
Veröffentlicht: (2024)
How to Train Your Fact Verifier: Knowledge Transfer with Multimodal Open Models
von: Lee, Jaeyoung, et al.
Veröffentlicht: (2024)
von: Lee, Jaeyoung, et al.
Veröffentlicht: (2024)
Measuring and Improving Attentiveness to Partial Inputs with Counterfactuals
von: Elazar, Yanai, et al.
Veröffentlicht: (2023)
von: Elazar, Yanai, et al.
Veröffentlicht: (2023)
Better Slow than Sorry: Introducing Positive Friction for Reliable Dialogue Systems
von: İnan, Mert, et al.
Veröffentlicht: (2025)
von: İnan, Mert, et al.
Veröffentlicht: (2025)
Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
Can Language Models Reason about Individualistic Human Values and Preferences?
von: Jiang, Liwei, et al.
Veröffentlicht: (2024)
von: Jiang, Liwei, et al.
Veröffentlicht: (2024)
RewardBench: Evaluating Reward Models for Language Modeling
von: Lambert, Nathan, et al.
Veröffentlicht: (2024)
von: Lambert, Nathan, et al.
Veröffentlicht: (2024)
Can Vision Language Models Understand Mimed Actions?
von: Cho, Hyundong, et al.
Veröffentlicht: (2025)
von: Cho, Hyundong, et al.
Veröffentlicht: (2025)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
von: Han, Seungju, et al.
Veröffentlicht: (2024)
von: Han, Seungju, et al.
Veröffentlicht: (2024)
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
Language Models can Infer Action Semantics for Symbolic Planners from Environment Feedback
von: Zhu, Wang, et al.
Veröffentlicht: (2024)
von: Zhu, Wang, et al.
Veröffentlicht: (2024)
Generating Contextually-Relevant Navigation Instructions for Blind and Low Vision People
von: Merchant, Zain, et al.
Veröffentlicht: (2024)
von: Merchant, Zain, et al.
Veröffentlicht: (2024)
CULTURE-GEN: Revealing Global Cultural Perception in Language Models through Natural Language Prompting
von: Li, Huihan, et al.
Veröffentlicht: (2024)
von: Li, Huihan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024) -
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
von: Chandu, Khyathi Raghavi, et al.
Veröffentlicht: (2024) -
L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects
von: Yamada, Yutaro, et al.
Veröffentlicht: (2024) -
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
von: Kamath, Amita, et al.
Veröffentlicht: (2026) -
Adjust for Trust: Mitigating Trust-Induced Inappropriate Reliance on AI Assistance
von: Srinivasan, Tejas, et al.
Veröffentlicht: (2025)