Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
Fuente:
arXiv
Saved in:
| Main Authors: | Chandu, Khyathi Raghavi, Li, Linjie, Awadalla, Anas, Lu, Ximing, Park, Jae Sung, Hessel, Jack, Wang, Lijuan, Choi, Yejin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Selective "Selective Prediction": Reducing Unnecessary Abstention in Vision-Language Reasoning
by: Srinivasan, Tejas, et al.
Published: (2024)
by: Srinivasan, Tejas, et al.
Published: (2024)
Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models
by: Sicilia, Anthony, et al.
Published: (2024)
by: Sicilia, Anthony, et al.
Published: (2024)
Synthetic Visual Genome
by: Park, Jae Sung, et al.
Published: (2025)
by: Park, Jae Sung, et al.
Published: (2025)
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
by: Kamath, Amita, et al.
Published: (2026)
by: Kamath, Amita, et al.
Published: (2026)
L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects
by: Yamada, Yutaro, et al.
Published: (2024)
by: Yamada, Yutaro, et al.
Published: (2024)
How to Train Your Fact Verifier: Knowledge Transfer with Multimodal Open Models
by: Lee, Jaeyoung, et al.
Published: (2024)
by: Lee, Jaeyoung, et al.
Published: (2024)
RESTOR: Knowledge Recovery in Machine Unlearning
by: Rezaei, Keivan, et al.
Published: (2024)
by: Rezaei, Keivan, et al.
Published: (2024)
AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
by: Lu, Ximing, et al.
Published: (2024)
by: Lu, Ximing, et al.
Published: (2024)
Continual Dialogue State Tracking via Example-Guided Question Answering
by: Cho, Hyundong, et al.
Published: (2023)
by: Cho, Hyundong, et al.
Published: (2023)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
by: Lin, Bill Yuchen, et al.
Published: (2024)
by: Lin, Bill Yuchen, et al.
Published: (2024)
The Art of Saying No: Contextual Noncompliance in Language Models
by: Brahman, Faeze, et al.
Published: (2024)
by: Brahman, Faeze, et al.
Published: (2024)
Tailoring Self-Rationalizers with Multi-Reward Distillation
by: Ramnath, Sahana, et al.
Published: (2023)
by: Ramnath, Sahana, et al.
Published: (2023)
Agent Lumos: Unified and Modular Training for Open-Source Language Agents
by: Yin, Da, et al.
Published: (2023)
by: Yin, Da, et al.
Published: (2023)
WildChat: 1M ChatGPT Interaction Logs in the Wild
by: Zhao, Wenting, et al.
Published: (2024)
by: Zhao, Wenting, et al.
Published: (2024)
WildVis: Open Source Visualizer for Million-Scale Chat Logs in the Wild
by: Deng, Yuntian, et al.
Published: (2024)
by: Deng, Yuntian, et al.
Published: (2024)
MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
by: Awadalla, Anas, et al.
Published: (2024)
by: Awadalla, Anas, et al.
Published: (2024)
Aleatoric and Epistemic Discrimination: Fundamental Limits of Fairness Interventions
by: Wang, Hao, et al.
Published: (2023)
by: Wang, Hao, et al.
Published: (2023)
Estimating Epistemic and Aleatoric Uncertainty with a Single Model
by: Chan, Matthew A., et al.
Published: (2024)
by: Chan, Matthew A., et al.
Published: (2024)
An Epistemic and Aleatoric Decomposition of Arbitrariness to Constrain the Set of Good Models
by: Khan, Falaah Arif, et al.
Published: (2023)
by: Khan, Falaah Arif, et al.
Published: (2023)
Symbolic Chain-of-Thought Distillation: Small Models Can Also "Think" Step-by-Step
by: Li, Liunian Harold, et al.
Published: (2023)
by: Li, Liunian Harold, et al.
Published: (2023)
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
by: Salehi, Mohammadreza, et al.
Published: (2024)
by: Salehi, Mohammadreza, et al.
Published: (2024)
Measuring Aleatoric and Epistemic Uncertainty in LLMs: Empirical Evaluation on ID and OOD QA Tasks
by: Wang, Kevin, et al.
Published: (2025)
by: Wang, Kevin, et al.
Published: (2025)
StyleRemix: Interpretable Authorship Obfuscation via Distillation and Perturbation of Style Elements
by: Fisher, Jillian, et al.
Published: (2024)
by: Fisher, Jillian, et al.
Published: (2024)
WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries
by: Zhao, Wenting, et al.
Published: (2024)
by: Zhao, Wenting, et al.
Published: (2024)
Calibrated Decomposition of Aleatoric and Epistemic Uncertainty in Deep Features for Inference-Time Adaptation
by: Kumar, Divake, et al.
Published: (2025)
by: Kumar, Divake, et al.
Published: (2025)
Measuring and Improving Attentiveness to Partial Inputs with Counterfactuals
by: Elazar, Yanai, et al.
Published: (2023)
by: Elazar, Yanai, et al.
Published: (2023)
JAMDEC: Unsupervised Authorship Obfuscation using Constrained Decoding over Small Language Models
by: Fisher, Jillian, et al.
Published: (2024)
by: Fisher, Jillian, et al.
Published: (2024)
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
by: Hao, Yunzhuo, et al.
Published: (2025)
by: Hao, Yunzhuo, et al.
Published: (2025)
Rethinking Epistemic and Aleatoric Uncertainty for Active Open-Set Annotation: An Energy-Based Approach
by: Zong, Chen-Chen, et al.
Published: (2025)
by: Zong, Chen-Chen, et al.
Published: (2025)
BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions
by: Awadalla, Anas, et al.
Published: (2024)
by: Awadalla, Anas, et al.
Published: (2024)
Bring Metric Functions into Diffusion Models
by: An, Jie, et al.
Published: (2024)
by: An, Jie, et al.
Published: (2024)
Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions
by: Acuna, David, et al.
Published: (2025)
by: Acuna, David, et al.
Published: (2025)
Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
by: Li, Linjie, et al.
Published: (2025)
by: Li, Linjie, et al.
Published: (2025)
Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
by: Cho, Jaemin, et al.
Published: (2023)
by: Cho, Jaemin, et al.
Published: (2023)
ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning
by: Gu, Jiawei, et al.
Published: (2025)
by: Gu, Jiawei, et al.
Published: (2025)
LUMA: A Benchmark Dataset for Learning from Uncertain and Multimodal Data
by: Bezirganyan, Grigor, et al.
Published: (2024)
by: Bezirganyan, Grigor, et al.
Published: (2024)
MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
by: Xia, Peng, et al.
Published: (2024)
by: Xia, Peng, et al.
Published: (2024)
MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities
by: Yu, Weihao, et al.
Published: (2024)
by: Yu, Weihao, et al.
Published: (2024)
UNcommonsense Reasoning: Abductive Reasoning about Uncommon Situations
by: Zhao, Wenting, et al.
Published: (2023)
by: Zhao, Wenting, et al.
Published: (2023)
Beyond Words: Advancing Long-Text Image Generation via Multimodal Autoregressive Models
by: Wang, Alex Jinpeng, et al.
Published: (2025)
by: Wang, Alex Jinpeng, et al.
Published: (2025)
Similar Items
-
Selective "Selective Prediction": Reducing Unnecessary Abstention in Vision-Language Reasoning
by: Srinivasan, Tejas, et al.
Published: (2024) -
Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models
by: Sicilia, Anthony, et al.
Published: (2024) -
Synthetic Visual Genome
by: Park, Jae Sung, et al.
Published: (2025) -
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
by: Kamath, Amita, et al.
Published: (2026) -
L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects
by: Yamada, Yutaro, et al.
Published: (2024)