UProp: Investigating the Uncertainty Propagation of LLMs in Multi-Step Agentic Decision-Making
Fuente:
arXiv
Saved in:
| Main Authors: | Duan, Jinhao, Diffenderfer, James, Madireddy, Sandeep, Chen, Tianlong, Kailkhura, Bhavya, Xu, Kaidi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations
by: Duan, Jinhao, et al.
Published: (2024)
by: Duan, Jinhao, et al.
Published: (2024)
TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-Intervention
by: Duan, Jinhao, et al.
Published: (2025)
by: Duan, Jinhao, et al.
Published: (2025)
Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models
by: Duan, Jinhao, et al.
Published: (2023)
by: Duan, Jinhao, et al.
Published: (2023)
LLM Unlearning Reveals a Stronger-Than-Expected Coreset Effect in Current Benchmarks
by: Pal, Soumyadeep, et al.
Published: (2025)
by: Pal, Soumyadeep, et al.
Published: (2025)
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
by: Hong, Junyuan, et al.
Published: (2024)
by: Hong, Junyuan, et al.
Published: (2024)
IUQ: Interrogative Uncertainty Quantification for Long-Form Large Language Model Generation
by: Fan, Haozhi, et al.
Published: (2026)
by: Fan, Haozhi, et al.
Published: (2026)
COIN: Uncertainty-Guarding Selective Question Answering for Foundation Models with Provable Risk Guarantees
by: Wang, Zhiyuan, et al.
Published: (2025)
by: Wang, Zhiyuan, et al.
Published: (2025)
Sparse Neurons Carry Strong Signals of Question Ambiguity in LLMs
by: Zhang, Zhuoxuan, et al.
Published: (2025)
by: Zhang, Zhuoxuan, et al.
Published: (2025)
Word-Sequence Entropy: Towards Uncertainty Estimation in Free-Form Medical Question Answering Applications and Beyond
by: Wang, Zhiyuan, et al.
Published: (2024)
by: Wang, Zhiyuan, et al.
Published: (2024)
Low-rank finetuning for LLMs: A fairness perspective
by: Das, Saswat, et al.
Published: (2024)
by: Das, Saswat, et al.
Published: (2024)
The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages
by: Onyame, Eric, et al.
Published: (2026)
by: Onyame, Eric, et al.
Published: (2026)
DynaCode: A Dynamic Complexity-Aware Code Benchmark for Evaluating Large Language Models in Code Generation
by: Hu, Wenhao, et al.
Published: (2025)
by: Hu, Wenhao, et al.
Published: (2025)
SConU: Selective Conformal Uncertainty in Large Language Models
by: Wang, Zhiyuan, et al.
Published: (2025)
by: Wang, Zhiyuan, et al.
Published: (2025)
Dialogue is Better Than Monologue: Instructing Medical LLMs via Strategical Conversations
by: Liu, Zijie, et al.
Published: (2025)
by: Liu, Zijie, et al.
Published: (2025)
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense
by: Ouyang, Yang, et al.
Published: (2025)
by: Ouyang, Yang, et al.
Published: (2025)
Tracking the Limits of Knowledge Propagation: How LLMs Fail at Multi-Step Reasoning with Conflicting Knowledge
by: Feng, Yiyang, et al.
Published: (2026)
by: Feng, Yiyang, et al.
Published: (2026)
ConU: Conformal Uncertainty in Large Language Models with Correctness Coverage Guarantees
by: Wang, Zhiyuan, et al.
Published: (2024)
by: Wang, Zhiyuan, et al.
Published: (2024)
STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
by: Wang, Zijun, et al.
Published: (2025)
by: Wang, Zijun, et al.
Published: (2025)
Cognitive Bias in Decision-Making with LLMs
by: Echterhoff, Jessica, et al.
Published: (2024)
by: Echterhoff, Jessica, et al.
Published: (2024)
Uncertainty as a Planning Signal: Multi-Turn Decision Making for Goal-Oriented Conversation
by: Ling, Xinyi, et al.
Published: (2026)
by: Ling, Xinyi, et al.
Published: (2026)
Extracting and Understanding the Superficial Knowledge in Alignment
by: Chen, Runjin, et al.
Published: (2025)
by: Chen, Runjin, et al.
Published: (2025)
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
by: Huang, Jen-tse, et al.
Published: (2024)
by: Huang, Jen-tse, et al.
Published: (2024)
Stephanie2: Thinking, Waiting, and Making Decisions Like Humans in Step-by-Step AI Social Chat
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
Mixture of Robust Experts (MoRE):A Robust Denoising Method towards multiple perturbations
by: Cheng, Hao, et al.
Published: (2021)
by: Cheng, Hao, et al.
Published: (2021)
Chance-constrained Flow Matching for High-Fidelity Constraint-aware Generation
by: Liang, Jinhao, et al.
Published: (2025)
by: Liang, Jinhao, et al.
Published: (2025)
Skin-in-the-Game: Decision Making via Multi-Stakeholder Alignment in LLMs
by: Sel, Bilgehan, et al.
Published: (2024)
by: Sel, Bilgehan, et al.
Published: (2024)
GuideLLM: Exploring LLM-Guided Conversation with Applications in Autobiography Interviewing
by: Duan, Jinhao, et al.
Published: (2025)
by: Duan, Jinhao, et al.
Published: (2025)
Answer, Refuse, or Guess? Investigating Risk-Aware Decision Making in Language Models
by: Wu, Cheng-Kuang, et al.
Published: (2025)
by: Wu, Cheng-Kuang, et al.
Published: (2025)
MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
by: Pandit, Shrey, et al.
Published: (2025)
by: Pandit, Shrey, et al.
Published: (2025)
MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-Making
by: Kim, Yubin, et al.
Published: (2024)
by: Kim, Yubin, et al.
Published: (2024)
Enhancing Decision-Making of Large Language Models via Actor-Critic
by: Dong, Heng, et al.
Published: (2025)
by: Dong, Heng, et al.
Published: (2025)
Agentic Uncertainty Quantification
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
Knowledge Tagging System on Math Questions via LLMs with Flexible Demonstration Retriever
by: Li, Hang, et al.
Published: (2024)
by: Li, Hang, et al.
Published: (2024)
LEC: Linear Expectation Constraints for Selection-Conditioned Risk Control in Selective Prediction and Routing Systems
by: Wang, Zhiyuan, et al.
Published: (2025)
by: Wang, Zhiyuan, et al.
Published: (2025)
Beyond Surface Statistics: Robust Conformal Prediction for LLMs via Internal Representations
by: Wang, Yanli, et al.
Published: (2026)
by: Wang, Yanli, et al.
Published: (2026)
SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese
by: Xu, Liang, et al.
Published: (2024)
by: Xu, Liang, et al.
Published: (2024)
DeLLMa: Decision Making Under Uncertainty with Large Language Models
by: Liu, Ollie, et al.
Published: (2024)
by: Liu, Ollie, et al.
Published: (2024)
Beyond MedQA: Towards Real-world Clinical Decision Making in the Era of LLMs
by: Xiao, Yunpeng, et al.
Published: (2025)
by: Xiao, Yunpeng, et al.
Published: (2025)
Evaluating LLMs for Police Decision-Making: A Framework Based on Police Action Scenarios
by: Lee, Sangyub, et al.
Published: (2026)
by: Lee, Sangyub, et al.
Published: (2026)
LLMs for Explainable Business Decision-Making: A Reinforcement Learning Fine-Tuning Approach
by: Cheng, Xiang, et al.
Published: (2025)
by: Cheng, Xiang, et al.
Published: (2025)
Similar Items
-
GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations
by: Duan, Jinhao, et al.
Published: (2024) -
TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-Intervention
by: Duan, Jinhao, et al.
Published: (2025) -
Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models
by: Duan, Jinhao, et al.
Published: (2023) -
LLM Unlearning Reveals a Stronger-Than-Expected Coreset Effect in Current Benchmarks
by: Pal, Soumyadeep, et al.
Published: (2025) -
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
by: Hong, Junyuan, et al.
Published: (2024)