Addressing Pitfalls in the Evaluation of Uncertainty Estimation Methods for Natural Language Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Ielanskyi, Mykyta, Schweighofer, Kajetan, Aichberger, Lukas, Hochreiter, Sepp |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Uncertainty Estimation through Semantically Diverse Language Generation
by: Aichberger, Lukas, et al.
Published: (2024)
by: Aichberger, Lukas, et al.
Published: (2024)
On Information-Theoretic Measures of Predictive Uncertainty
by: Schweighofer, Kajetan, et al.
Published: (2024)
by: Schweighofer, Kajetan, et al.
Published: (2024)
Rethinking Uncertainty Estimation in LLMs: A Principled Single-Sequence Measure
by: Aichberger, Lukas, et al.
Published: (2024)
by: Aichberger, Lukas, et al.
Published: (2024)
The Disparate Benefits of Deep Ensembles
by: Schweighofer, Kajetan, et al.
Published: (2024)
by: Schweighofer, Kajetan, et al.
Published: (2024)
Unlocking the Working Memory of Large Language Models for Latent Reasoning
by: Aichberger, Lukas, et al.
Published: (2026)
by: Aichberger, Lukas, et al.
Published: (2026)
Uncertainty Quantification for Regression using Proper Scoring Rules
by: Fishkov, Alexander, et al.
Published: (2025)
by: Fishkov, Alexander, et al.
Published: (2025)
xLSTM Scaling Laws: Competitive Performance with Linear Time-Complexity
by: Beck, Maximilian, et al.
Published: (2025)
by: Beck, Maximilian, et al.
Published: (2025)
MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular Graphs
by: Bartmann, Christoph, et al.
Published: (2026)
by: Bartmann, Christoph, et al.
Published: (2026)
Overcoming Forgetting in LLM Fine-Tuning with Evolution Strategies
by: Schweighofer, Kajetan, et al.
Published: (2026)
by: Schweighofer, Kajetan, et al.
Published: (2026)
FlashRNN: I/O-Aware Optimization of Traditional RNNs on modern hardware
by: Pöppel, Korbinian, et al.
Published: (2024)
by: Pöppel, Korbinian, et al.
Published: (2024)
Efficient Pre-Training of LLMs through Truncated SVD Layers
by: Kamali, Kaivan, et al.
Published: (2026)
by: Kamali, Kaivan, et al.
Published: (2026)
Rethinking Losses for Diffusion Bridge Samplers
by: Sanokowski, Sebastian, et al.
Published: (2025)
by: Sanokowski, Sebastian, et al.
Published: (2025)
Safe and Certifiable AI Systems: Concepts, Challenges, and Lessons Learned
by: Schweighofer, Kajetan, et al.
Published: (2025)
by: Schweighofer, Kajetan, et al.
Published: (2025)
A Diffusion Model Framework for Unsupervised Neural Combinatorial Optimization
by: Sanokowski, Sebastian, et al.
Published: (2024)
by: Sanokowski, Sebastian, et al.
Published: (2024)
MIM-Refiner: A Contrastive Learning Boost from Intermediate Pre-Trained Representations
by: Alkin, Benedikt, et al.
Published: (2024)
by: Alkin, Benedikt, et al.
Published: (2024)
Tiled Flash Linear Attention: More Efficient Linear RNN and xLSTM Kernels
by: Beck, Maximilian, et al.
Published: (2025)
by: Beck, Maximilian, et al.
Published: (2025)
Contrastive Abstraction for Reinforcement Learning
by: Patil, Vihang, et al.
Published: (2024)
by: Patil, Vihang, et al.
Published: (2024)
Parameter Efficient Fine-tuning via Explained Variance Adaptation
by: Paischer, Fabian, et al.
Published: (2024)
by: Paischer, Fabian, et al.
Published: (2024)
Vision-LSTM: xLSTM as Generic Vision Backbone
by: Alkin, Benedikt, et al.
Published: (2024)
by: Alkin, Benedikt, et al.
Published: (2024)
Retrieval-Augmented Decision Transformer: External Memory for In-context RL
by: Schmied, Thomas, et al.
Published: (2024)
by: Schmied, Thomas, et al.
Published: (2024)
Pitfalls in Evaluating Language Model Forecasters
by: Paleka, Daniel, et al.
Published: (2025)
by: Paleka, Daniel, et al.
Published: (2025)
VN-EGNN: E(3)-Equivariant Graph Neural Networks with Virtual Nodes Enhance Protein Binding Site Identification
by: Sestak, Florian, et al.
Published: (2024)
by: Sestak, Florian, et al.
Published: (2024)
Large Language Models Can Self-Improve At Web Agent Tasks
by: Patel, Ajay, et al.
Published: (2024)
by: Patel, Ajay, et al.
Published: (2024)
Reconsidering LLM Uncertainty Estimation Methods in the Wild
by: Bakman, Yavuz, et al.
Published: (2025)
by: Bakman, Yavuz, et al.
Published: (2025)
The Pitfalls of Memorization: When Memorization Hurts Generalization
by: Bayat, Reza, et al.
Published: (2024)
by: Bayat, Reza, et al.
Published: (2024)
On Subjective Uncertainty Quantification and Calibration in Natural Language Generation
by: Wang, Ziyu, et al.
Published: (2024)
by: Wang, Ziyu, et al.
Published: (2024)
From Entropy to Calibrated Uncertainty: Training Language Models to Reason About Uncertainty
by: Jenane, Azza, et al.
Published: (2026)
by: Jenane, Azza, et al.
Published: (2026)
A Large Recurrent Action Model: xLSTM enables Fast Inference for Robotics Tasks
by: Schmied, Thomas, et al.
Published: (2024)
by: Schmied, Thomas, et al.
Published: (2024)
xLSTM: Extended Long Short-Term Memory
by: Beck, Maximilian, et al.
Published: (2024)
by: Beck, Maximilian, et al.
Published: (2024)
Bio-xLSTM: Generative modeling, representation and in-context learning of biological and chemical sequences
by: Schmidinger, Niklas, et al.
Published: (2024)
by: Schmidinger, Niklas, et al.
Published: (2024)
Efficient Nearest Neighbor based Uncertainty Estimation for Natural Language Processing Tasks
by: Hashimoto, Wataru, et al.
Published: (2024)
by: Hashimoto, Wataru, et al.
Published: (2024)
On Uncertainty In Natural Language Processing
by: Ulmer, Dennis
Published: (2024)
by: Ulmer, Dennis
Published: (2024)
SymbolicAI: A framework for logic-based approaches combining generative models and solvers
by: Dinu, Marius-Constantin, et al.
Published: (2024)
by: Dinu, Marius-Constantin, et al.
Published: (2024)
Scalable and Efficient Methods for Uncertainty Estimation and Reduction in Deep Learning
by: Ahmed, Soyed Tuhin
Published: (2024)
by: Ahmed, Soyed Tuhin
Published: (2024)
Evaluating Supervised Machine Learning Models: Principles, Pitfalls, and Metric Selection
by: Liu, Xuanyan, et al.
Published: (2026)
by: Liu, Xuanyan, et al.
Published: (2026)
A Generic Method for Fine-grained Category Discovery in Natural Language Texts
by: Tian, Chang, et al.
Published: (2024)
by: Tian, Chang, et al.
Published: (2024)
The Pitfalls of KV Cache Compression
by: Chen, Alex, et al.
Published: (2025)
by: Chen, Alex, et al.
Published: (2025)
General Uncertainty Estimation with Delta Variances
by: Schmitt, Simon, et al.
Published: (2025)
by: Schmitt, Simon, et al.
Published: (2025)
Scalable Discrete Diffusion Samplers: Combinatorial Optimization and Statistical Physics
by: Sanokowski, Sebastian, et al.
Published: (2025)
by: Sanokowski, Sebastian, et al.
Published: (2025)
xLSTM 7B: A Recurrent LLM for Fast and Efficient Inference
by: Beck, Maximilian, et al.
Published: (2025)
by: Beck, Maximilian, et al.
Published: (2025)
Similar Items
-
Improving Uncertainty Estimation through Semantically Diverse Language Generation
by: Aichberger, Lukas, et al.
Published: (2024) -
On Information-Theoretic Measures of Predictive Uncertainty
by: Schweighofer, Kajetan, et al.
Published: (2024) -
Rethinking Uncertainty Estimation in LLMs: A Principled Single-Sequence Measure
by: Aichberger, Lukas, et al.
Published: (2024) -
The Disparate Benefits of Deep Ensembles
by: Schweighofer, Kajetan, et al.
Published: (2024) -
Unlocking the Working Memory of Large Language Models for Latent Reasoning
by: Aichberger, Lukas, et al.
Published: (2026)