Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results
Fuente:
arXiv
Saved in:
| Main Authors: | Santilli, Andrea, Golinski, Adam, Kirchhof, Michael, Danieli, Federico, Blaas, Arno, Xiong, Miao, Zappella, Luca, Williamson, Sinead |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Uncertainty Quantification for LLM Function-Calling
by: Ye, Zihuiwen, et al.
Published: (2026)
by: Ye, Zihuiwen, et al.
Published: (2026)
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
by: Kirchhof, Michael, et al.
Published: (2025)
by: Kirchhof, Michael, et al.
Published: (2025)
Considerations for Distribution Shift Robustness of Diagnostic Models in Healthcare
by: Blaas, Arno, et al.
Published: (2024)
by: Blaas, Arno, et al.
Published: (2024)
GenCtrl -- A Formal Controllability Toolkit for Generative Models
by: Cheng, Emily, et al.
Published: (2026)
by: Cheng, Emily, et al.
Published: (2026)
Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs
by: Nakkiran, Preetum, et al.
Published: (2025)
by: Nakkiran, Preetum, et al.
Published: (2025)
Interpreting CLIP: Insights on the Robustness to ImageNet Distribution Shifts
by: Crabbé, Jonathan, et al.
Published: (2023)
by: Crabbé, Jonathan, et al.
Published: (2023)
BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design
by: Choudhury, Deepro, et al.
Published: (2025)
by: Choudhury, Deepro, et al.
Published: (2025)
Posterior Uncertainty Quantification in Neural Networks using Data Augmentation
by: Wu, Luhuan, et al.
Published: (2024)
by: Wu, Luhuan, et al.
Published: (2024)
Controlling Language and Diffusion Models by Transporting Activations
by: Rodriguez, Pau, et al.
Published: (2024)
by: Rodriguez, Pau, et al.
Published: (2024)
ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language Models
by: Danieli, Federico, et al.
Published: (2025)
by: Danieli, Federico, et al.
Published: (2025)
Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs
by: Wang, Yinong Oliver, et al.
Published: (2025)
by: Wang, Yinong Oliver, et al.
Published: (2025)
LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss
by: Rodriguez, Pau, et al.
Published: (2025)
by: Rodriguez, Pau, et al.
Published: (2025)
Empirical Quantification of Spurious Correlations in Malware Detection
by: Perasso, Bianca, et al.
Published: (2025)
by: Perasso, Bianca, et al.
Published: (2025)
Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents
by: Kirchhof, Michael, et al.
Published: (2025)
by: Kirchhof, Michael, et al.
Published: (2025)
Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity
by: Huang, Ningyuan, et al.
Published: (2025)
by: Huang, Ningyuan, et al.
Published: (2025)
Trace Length is a Simple Uncertainty Signal in Reasoning Models
by: Devic, Siddartha, et al.
Published: (2025)
by: Devic, Siddartha, et al.
Published: (2025)
LLMs are not (consistently) Bayesian: Quantifying internal (in)consistencies of LLMs' probabilistic beliefs
by: Chen, Chacha, et al.
Published: (2026)
by: Chen, Chacha, et al.
Published: (2026)
Uncertainties of Latent Representations in Computer Vision
by: Kirchhof, Michael
Published: (2024)
by: Kirchhof, Michael
Published: (2024)
Attention to Mamba: A Recipe for Cross-Architecture Distillation
by: Moudgil, Abhinav, et al.
Published: (2026)
by: Moudgil, Abhinav, et al.
Published: (2026)
Uncertainty Quantification for Evaluating Machine Translation Bias
by: Staliūnaitė, Ieva Raminta, et al.
Published: (2025)
by: Staliūnaitė, Ieva Raminta, et al.
Published: (2025)
Self-Supervised Learning with Gaussian Processes
by: Duan, Yunshan, et al.
Published: (2025)
by: Duan, Yunshan, et al.
Published: (2025)
Bias in Gender Bias Benchmarks: How Spurious Features Distort Evaluation
by: Hirota, Yusuke, et al.
Published: (2025)
by: Hirota, Yusuke, et al.
Published: (2025)
Benchmarking Uncertainty Disentanglement: Specialized Uncertainties for Specialized Tasks
by: Mucsányi, Bálint, et al.
Published: (2024)
by: Mucsányi, Bálint, et al.
Published: (2024)
Evaluating Gender Bias Transfer between Pre-trained and Prompt-Adapted Language Models
by: Mackraz, Natalie, et al.
Published: (2024)
by: Mackraz, Natalie, et al.
Published: (2024)
A note on precotangent bundles: the example of Grassmannians
by: Goliński, Tomasz
Published: (2025)
by: Goliński, Tomasz
Published: (2025)
Deep Gaussian Process Emulation and Uncertainty Quantification for Large Computer Experiments
by: Yazdi, Faezeh, et al.
Published: (2024)
by: Yazdi, Faezeh, et al.
Published: (2024)
Revisiting Spurious Correlation in Domain Generalization
by: Qin, Bin, et al.
Published: (2024)
by: Qin, Bin, et al.
Published: (2024)
Pretrained Visual Uncertainties
by: Kirchhof, Michael, et al.
Published: (2024)
by: Kirchhof, Michael, et al.
Published: (2024)
Explaining Length Bias in LLM-Based Preference Evaluations
by: Hu, Zhengyu, et al.
Published: (2024)
by: Hu, Zhengyu, et al.
Published: (2024)
Bias after Prompting: Persistent Discrimination in Large Language Models
by: Sivakumar, Nivedha, et al.
Published: (2025)
by: Sivakumar, Nivedha, et al.
Published: (2025)
Agentic Uncertainty Quantification
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
Uncertainty Propagation in Stochastic Systems via Mixture Models with Error Quantification
by: Figueiredo, Eduardo, et al.
Published: (2024)
by: Figueiredo, Eduardo, et al.
Published: (2024)
To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models
by: Malach, Eran, et al.
Published: (2025)
by: Malach, Eran, et al.
Published: (2025)
DSO: Direct Steering Optimization for Bias Mitigation
by: Paes, Lucas Monteiro, et al.
Published: (2025)
by: Paes, Lucas Monteiro, et al.
Published: (2025)
Evaluating Uncertainty Quantification Methods in Argumentative Large Language Models
by: Zhou, Kevin, et al.
Published: (2025)
by: Zhou, Kevin, et al.
Published: (2025)
Análise da importância da utilização do orçamento e do planejamento estratégico como ferramenta de controle na atividade rural
by: Lizandra Blaas dos Santos
Published: (2011)
by: Lizandra Blaas dos Santos
Published: (2011)
Fast & fuelious: the malate–aspartate shuttle in brown adipocyte lipid metabolism
by: Lukas Blaas, et al.
Published: (2026)
by: Lukas Blaas, et al.
Published: (2026)
SAF‐guarding the cuff: Could shoulder fat cells combat fibrosis?
by: Lukas Blaas, et al.
Published: (2025)
by: Lukas Blaas, et al.
Published: (2025)
Sparse Autoencoders are Capable LLM Jailbreak Mitigators
by: Assogba, Yannick, et al.
Published: (2026)
by: Assogba, Yannick, et al.
Published: (2026)
Pesquisa colaborativa: refl exões sobre o processo de aprendizagem como prática social de uma professora pré-serviço de inglês
by: Luciane Kirchhof Ticks
Published: (2011)
by: Luciane Kirchhof Ticks
Published: (2011)
Similar Items
-
Uncertainty Quantification for LLM Function-Calling
by: Ye, Zihuiwen, et al.
Published: (2026) -
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
by: Kirchhof, Michael, et al.
Published: (2025) -
Considerations for Distribution Shift Robustness of Diagnostic Models in Healthcare
by: Blaas, Arno, et al.
Published: (2024) -
GenCtrl -- A Formal Controllability Toolkit for Generative Models
by: Cheng, Emily, et al.
Published: (2026) -
Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs
by: Nakkiran, Preetum, et al.
Published: (2025)