Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Kirchhof, Michael, Kasneci, Gjergji, Kasneci, Enkelejda |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enriching Tabular Data with Contextual LLM Embeddings: A Comprehensive Ablation Study for Ensemble Classifiers
by: Kasneci, Gjergji, et al.
Published: (2024)
by: Kasneci, Gjergji, et al.
Published: (2024)
Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
by: Kasneci, Enkelejda, et al.
Published: (2026)
by: Kasneci, Enkelejda, et al.
Published: (2026)
Emergent Abilities in Large Language Models: A Survey
by: Berti, Leonardo, et al.
Published: (2025)
by: Berti, Leonardo, et al.
Published: (2025)
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
by: Leemann, Tobias, et al.
Published: (2024)
by: Leemann, Tobias, et al.
Published: (2024)
CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models
by: Kocak, Aysenur, et al.
Published: (2025)
by: Kocak, Aysenur, et al.
Published: (2025)
Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods
by: Dhaini, Mahdi, et al.
Published: (2025)
by: Dhaini, Mahdi, et al.
Published: (2025)
EvalxNLP: A Framework for Benchmarking Post-Hoc Explainability Methods on NLP Models
by: Dhaini, Mahdi, et al.
Published: (2025)
by: Dhaini, Mahdi, et al.
Published: (2025)
Stepwise Self-Consistent Mathematical Reasoning with Large Language Models
by: Zhao, Zilong, et al.
Published: (2024)
by: Zhao, Zilong, et al.
Published: (2024)
Is Crowdsourcing Breaking Your Bank? Cost-Effective Fine-Tuning of Pre-trained Language Models with Proximal Policy Optimization
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
Multimodal Behavioral Patterns Analysis with Eye-Tracking and LLM-Based Reasoning
by: Guo, Dongyang, et al.
Published: (2025)
by: Guo, Dongyang, et al.
Published: (2025)
Understanding Knowledge Drift in LLMs through Misinformation
by: Fastowski, Alina, et al.
Published: (2024)
by: Fastowski, Alina, et al.
Published: (2024)
Where Paths Split: Localized, Calibrated Control of Moral Reasoning in Large Language Models
by: Yuan, Chenchen, et al.
Published: (2026)
by: Yuan, Chenchen, et al.
Published: (2026)
Pretrained Visual Uncertainties
by: Kirchhof, Michael, et al.
Published: (2024)
by: Kirchhof, Michael, et al.
Published: (2024)
TLOB: A Novel Transformer Model with Dual Attention for Price Trend Prediction with Limit Order Book Data
by: Berti, Leonardo, et al.
Published: (2025)
by: Berti, Leonardo, et al.
Published: (2025)
From Confidence to Collapse in LLM Factual Robustness
by: Fastowski, Alina, et al.
Published: (2025)
by: Fastowski, Alina, et al.
Published: (2025)
Grokking in the Wild: Data Augmentation for Real-World Multi-Hop Reasoning with Transformers
by: Abramov, Roman, et al.
Published: (2025)
by: Abramov, Roman, et al.
Published: (2025)
MemeScouts@LT-EDI 2026: Asking the Right Questions -- Prompted Weak Supervision for Meme Hate Speech Detection
by: Bueno, Ivo, et al.
Published: (2026)
by: Bueno, Ivo, et al.
Published: (2026)
Benchmarking Large Language Models for Math Reasoning Tasks
by: Seßler, Kathrin, et al.
Published: (2024)
by: Seßler, Kathrin, et al.
Published: (2024)
Consolidating Rewarded Perturbations for LLM Post-Training
by: Zhang, Zheyu, et al.
Published: (2026)
by: Zhang, Zheyu, et al.
Published: (2026)
RAZOR: Sharpening Knowledge by Cutting Bias with Unsupervised Text Rewriting
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
I Prefer not to Say: Protecting User Consent in Models with Optional Personal Data
by: Leemann, Tobias, et al.
Published: (2022)
by: Leemann, Tobias, et al.
Published: (2022)
Multimodal Assessment of Classroom Discourse Quality: A Text-Centered Attention-Based Multi-Task Learning Approach
by: Hou, Ruikun, et al.
Published: (2025)
by: Hou, Ruikun, et al.
Published: (2025)
Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring
by: Seßler, Kathrin, et al.
Published: (2024)
by: Seßler, Kathrin, et al.
Published: (2024)
Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results
by: Santilli, Andrea, et al.
Published: (2025)
by: Santilli, Andrea, et al.
Published: (2025)
Probabilistic Aggregation and Targeted Embedding Optimization for Collective Moral Reasoning in Large Language Models
by: Yuan, Chenchen, et al.
Published: (2025)
by: Yuan, Chenchen, et al.
Published: (2025)
Faithful Attention Explainer: Verbalizing Decisions Based on Discriminative Features
by: Rong, Yao, et al.
Published: (2024)
by: Rong, Yao, et al.
Published: (2024)
Exploring User Acceptance and Concerns toward LLM-powered Conversational Agents in Immersive Extended Reality
by: Bozkir, Efe, et al.
Published: (2025)
by: Bozkir, Efe, et al.
Published: (2025)
Not All Features Deserve Attention: Graph-Guided Dependency Learning for Tabular Data Generation with Language Models
by: Zhang, Zheyu, et al.
Published: (2025)
by: Zhang, Zheyu, et al.
Published: (2025)
I-CEE: Tailoring Explanations of Image Classification Models to User Expertise
by: Rong, Yao, et al.
Published: (2023)
by: Rong, Yao, et al.
Published: (2023)
Doubling Your Data in Minutes: Ultra-fast Tabular Data Generation via LLM-Induced Dependency Graphs
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
Active Tabular Augmentation via Policy-Guided Diffusion Inpainting
by: Zhang, Zheyu, et al.
Published: (2026)
by: Zhang, Zheyu, et al.
Published: (2026)
Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs
by: Yuan, Chenchen, et al.
Published: (2026)
by: Yuan, Chenchen, et al.
Published: (2026)
Can LLM-Generated Textual Explanations Enhance Model Classification Performance? An Empirical Study
by: Dhaini, Mahdi, et al.
Published: (2025)
by: Dhaini, Mahdi, et al.
Published: (2025)
The Consistency Hypothesis in Uncertainty Quantification for Large Language Models
by: Xiao, Quan, et al.
Published: (2025)
by: Xiao, Quan, et al.
Published: (2025)
Adoption of Explainable Natural Language Processing: Perspectives from Industry and Academia on Practices and Challenges
by: Dhaini, Mahdi, et al.
Published: (2025)
by: Dhaini, Mahdi, et al.
Published: (2025)
Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs
by: Fastowski, Alina, et al.
Published: (2025)
by: Fastowski, Alina, et al.
Published: (2025)
TEyeD: Over 20 million real-world eye images with Pupil, Eyelid, and Iris 2D and 3D Segmentations, 2D and 3D Landmarks, 3D Eyeball, Gaze Vector, and Eye Movement Types
by: Fuhl, Wolfgang, et al.
Published: (2021)
by: Fuhl, Wolfgang, et al.
Published: (2021)
Taking the Next Step with Generative Artificial Intelligence: The Transformative Role of Multimodal Large Language Models in Science Education
by: Bewersdorff, Arne, et al.
Published: (2024)
by: Bewersdorff, Arne, et al.
Published: (2024)
Assessing the Real-World Utility of Explainable AI for Arousal Diagnostics: An Application-Grounded User Study
by: Kraft, Stefan, et al.
Published: (2025)
by: Kraft, Stefan, et al.
Published: (2025)
Entry Dependent Expert Selection in Distributed Gaussian Processes Using Multilabel Classification
by: Jalali, Hamed, et al.
Published: (2022)
by: Jalali, Hamed, et al.
Published: (2022)
Similar Items
-
Enriching Tabular Data with Contextual LLM Embeddings: A Comprehensive Ablation Study for Ensemble Classifiers
by: Kasneci, Gjergji, et al.
Published: (2024) -
Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
by: Kasneci, Enkelejda, et al.
Published: (2026) -
Emergent Abilities in Large Language Models: A Survey
by: Berti, Leonardo, et al.
Published: (2025) -
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
by: Leemann, Tobias, et al.
Published: (2024) -
CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models
by: Kocak, Aysenur, et al.
Published: (2025)