To Predict or Not to Predict? Towards reliable uncertainty estimation in the presence of noise
Fuente:
arXiv
Saved in:
| Main Authors: | Khallaf, Nouran, Sharoff, Serge |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Much Noise Can BERT Handle? Insights from Multilingual Sentence Difficulty Detection
by: Khallaf, Nouran, et al.
Published: (2026)
by: Khallaf, Nouran, et al.
Published: (2026)
Reading Between the Lines: A dataset and a study on why some texts are tougher than others
by: Khallaf, Nouran, et al.
Published: (2025)
by: Khallaf, Nouran, et al.
Published: (2025)
Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification
by: Hilasaca, Kenji, et al.
Published: (2026)
by: Hilasaca, Kenji, et al.
Published: (2026)
Can LLM Reasoning Be Trusted? A Comparative Study: Using Human Benchmarking on Statistical Tasks
by: Nagarkar, Crish, et al.
Published: (2026)
by: Nagarkar, Crish, et al.
Published: (2026)
Controlling Out-of-Domain Gaps in LLMs for Genre Classification and Generated Text Detection
by: Roussinov, Dmitri, et al.
Published: (2024)
by: Roussinov, Dmitri, et al.
Published: (2024)
A Multilingual Human Annotated Corpus of Original and Easy-to-Read Texts to Support Access to Democratic Participatory Processes
by: Bott, Stefan, et al.
Published: (2026)
by: Bott, Stefan, et al.
Published: (2026)
Almost Clinical: Linguistic properties of synthetic electronic health records
by: Sharoff, Serge, et al.
Published: (2026)
by: Sharoff, Serge, et al.
Published: (2026)
UoL-UPF at TSAR 2025 Shared Task A Generate-and-Select Approach for Readability-Controlled Text Simplification.
by: Hayakawa, Akio, et al.
Published: (2025)
by: Hayakawa, Akio, et al.
Published: (2025)
Predict the Next Word: Humans exhibit uncertainty in this task and language models _____
by: Ilia, Evgenia, et al.
Published: (2024)
by: Ilia, Evgenia, et al.
Published: (2024)
Entropy trajectory shape predicts LLM reasoning reliability: A diagnostic study of uncertainty dynamics in chain-of-thought
by: Zhao, Xinghao
Published: (2026)
by: Zhao, Xinghao
Published: (2026)
Do LLMs estimate uncertainty well in instruction-following?
by: Heo, Juyeon, et al.
Published: (2024)
by: Heo, Juyeon, et al.
Published: (2024)
The Prediction-Measurement Gap: Toward Meaning Representations as Scientific Instruments
by: Plisiecki, Hubert
Published: (2026)
by: Plisiecki, Hubert
Published: (2026)
Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation
by: Carandang, Kristine Ann M., et al.
Published: (2025)
by: Carandang, Kristine Ann M., et al.
Published: (2025)
Towards Explainability in Legal Outcome Prediction Models
by: Valvoda, Josef, et al.
Published: (2024)
by: Valvoda, Josef, et al.
Published: (2024)
Towards a Diagnostic and Predictive Evaluation Methodology for Sequence Labeling Tasks
by: Alvarez-Mellado, Elena, et al.
Published: (2026)
by: Alvarez-Mellado, Elena, et al.
Published: (2026)
The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination
by: Zhang, Yuji, et al.
Published: (2025)
by: Zhang, Yuji, et al.
Published: (2025)
Prediction-powered estimators for finite population statistics in highly imbalanced textual data: Public hate crime estimation
by: Waldetoft, Hannes, et al.
Published: (2025)
by: Waldetoft, Hannes, et al.
Published: (2025)
Predictive Data Selection: The Data That Predicts Is the Data That Teaches
by: Shum, Kashun, et al.
Published: (2025)
by: Shum, Kashun, et al.
Published: (2025)
Predicting Through Generation: Why Generation Is Better for Prediction
by: Kowsher, Md, et al.
Published: (2025)
by: Kowsher, Md, et al.
Published: (2025)
Towards Predictive Communication with Brain-Computer Interfaces integrating Large Language Models
by: Caria, Andrea
Published: (2024)
by: Caria, Andrea
Published: (2024)
Towards Predicting Any Human Trajectory In Context
by: Fujii, Ryo, et al.
Published: (2025)
by: Fujii, Ryo, et al.
Published: (2025)
DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
by: Habba, Eliya, et al.
Published: (2025)
by: Habba, Eliya, et al.
Published: (2025)
Continuous Risk Prediction
by: Dai, Yi
Published: (2024)
by: Dai, Yi
Published: (2024)
Towards Reducing Diagnostic Errors with Interpretable Risk Prediction
by: McInerney, Denis Jered, et al.
Published: (2024)
by: McInerney, Denis Jered, et al.
Published: (2024)
Towards a Psychology of Machines: Large Language Models Predict Human Memory
by: Huff, Markus, et al.
Published: (2024)
by: Huff, Markus, et al.
Published: (2024)
Towards Reliable and Interpretable Traffic Crash Pattern Prediction and Safety Interventions Using Customized Large Language Models
by: Zhao, Yang, et al.
Published: (2025)
by: Zhao, Yang, et al.
Published: (2025)
Towards Robust and Fair Next Visit Diagnosis Prediction under Noisy Clinical Notes with Large Language Models
by: Koo, Heejoon
Published: (2025)
by: Koo, Heejoon
Published: (2025)
Beyond the Next Token: Towards Prompt-Robust Zero-Shot Classification via Efficient Multi-Token Prediction
by: Qian, Junlang, et al.
Published: (2025)
by: Qian, Junlang, et al.
Published: (2025)
Towards Better Graph-based Cross-document Relation Extraction via Non-bridge Entity Enhancement and Prediction Debiasing
by: Yue, Hao, et al.
Published: (2024)
by: Yue, Hao, et al.
Published: (2024)
The Craft of Selective Prediction: Towards Reliable Case Outcome Classification -- An Empirical Study on European Court of Human Rights Cases
by: Santosh, T. Y. S. S., et al.
Published: (2024)
by: Santosh, T. Y. S. S., et al.
Published: (2024)
Towards Auto-Regressive Next-Token Prediction: In-Context Learning Emerges from Generalization
by: Gong, Zixuan, et al.
Published: (2025)
by: Gong, Zixuan, et al.
Published: (2025)
Beyond prompt brittleness: Evaluating the reliability and consistency of political worldviews in LLMs
by: Ceron, Tanise, et al.
Published: (2024)
by: Ceron, Tanise, et al.
Published: (2024)
Legal Fact Prediction: The Missing Piece in Legal Judgment Prediction
by: Liu, Junkai, et al.
Published: (2024)
by: Liu, Junkai, et al.
Published: (2024)
Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining
by: Peng, Bo, et al.
Published: (2026)
by: Peng, Bo, et al.
Published: (2026)
Non-Linear Scoring Model for Translation Quality Evaluation
by: Gladkoff, Serge, et al.
Published: (2025)
by: Gladkoff, Serge, et al.
Published: (2025)
Exploring Fact Memorization and Style Imitation in LLMs Using QLoRA: An Experimental Study and Quality Assessment Methods
by: Vyborov, Eugene, et al.
Published: (2024)
by: Vyborov, Eugene, et al.
Published: (2024)
Ideology Prediction of German Political Texts
by: Schneider, Sinclair, et al.
Published: (2026)
by: Schneider, Sinclair, et al.
Published: (2026)
Reward Prediction with Factorized World States
by: Shen, Yijun, et al.
Published: (2026)
by: Shen, Yijun, et al.
Published: (2026)
Evaluating the Homogeneity of Keyphrase Prediction Models
by: Houbre, Maël, et al.
Published: (2026)
by: Houbre, Maël, et al.
Published: (2026)
Promptly Predicting Structures: The Return of Inference
by: Mehta, Maitrey, et al.
Published: (2024)
by: Mehta, Maitrey, et al.
Published: (2024)
Similar Items
-
How Much Noise Can BERT Handle? Insights from Multilingual Sentence Difficulty Detection
by: Khallaf, Nouran, et al.
Published: (2026) -
Reading Between the Lines: A dataset and a study on why some texts are tougher than others
by: Khallaf, Nouran, et al.
Published: (2025) -
Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification
by: Hilasaca, Kenji, et al.
Published: (2026) -
Can LLM Reasoning Be Trusted? A Comparative Study: Using Human Benchmarking on Statistical Tasks
by: Nagarkar, Crish, et al.
Published: (2026) -
Controlling Out-of-Domain Gaps in LLMs for Genre Classification and Generated Text Detection
by: Roussinov, Dmitri, et al.
Published: (2024)