To Predict or Not to Predict? Towards reliable uncertainty estimation in the presence of noise
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Khallaf, Nouran, Sharoff, Serge |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Much Noise Can BERT Handle? Insights from Multilingual Sentence Difficulty Detection
von: Khallaf, Nouran, et al.
Veröffentlicht: (2026)
von: Khallaf, Nouran, et al.
Veröffentlicht: (2026)
Reading Between the Lines: A dataset and a study on why some texts are tougher than others
von: Khallaf, Nouran, et al.
Veröffentlicht: (2025)
von: Khallaf, Nouran, et al.
Veröffentlicht: (2025)
Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification
von: Hilasaca, Kenji, et al.
Veröffentlicht: (2026)
von: Hilasaca, Kenji, et al.
Veröffentlicht: (2026)
Can LLM Reasoning Be Trusted? A Comparative Study: Using Human Benchmarking on Statistical Tasks
von: Nagarkar, Crish, et al.
Veröffentlicht: (2026)
von: Nagarkar, Crish, et al.
Veröffentlicht: (2026)
Controlling Out-of-Domain Gaps in LLMs for Genre Classification and Generated Text Detection
von: Roussinov, Dmitri, et al.
Veröffentlicht: (2024)
von: Roussinov, Dmitri, et al.
Veröffentlicht: (2024)
A Multilingual Human Annotated Corpus of Original and Easy-to-Read Texts to Support Access to Democratic Participatory Processes
von: Bott, Stefan, et al.
Veröffentlicht: (2026)
von: Bott, Stefan, et al.
Veröffentlicht: (2026)
Almost Clinical: Linguistic properties of synthetic electronic health records
von: Sharoff, Serge, et al.
Veröffentlicht: (2026)
von: Sharoff, Serge, et al.
Veröffentlicht: (2026)
UoL-UPF at TSAR 2025 Shared Task A Generate-and-Select Approach for Readability-Controlled Text Simplification.
von: Hayakawa, Akio, et al.
Veröffentlicht: (2025)
von: Hayakawa, Akio, et al.
Veröffentlicht: (2025)
Predict the Next Word: Humans exhibit uncertainty in this task and language models _____
von: Ilia, Evgenia, et al.
Veröffentlicht: (2024)
von: Ilia, Evgenia, et al.
Veröffentlicht: (2024)
Entropy trajectory shape predicts LLM reasoning reliability: A diagnostic study of uncertainty dynamics in chain-of-thought
von: Zhao, Xinghao
Veröffentlicht: (2026)
von: Zhao, Xinghao
Veröffentlicht: (2026)
Do LLMs estimate uncertainty well in instruction-following?
von: Heo, Juyeon, et al.
Veröffentlicht: (2024)
von: Heo, Juyeon, et al.
Veröffentlicht: (2024)
The Prediction-Measurement Gap: Toward Meaning Representations as Scientific Instruments
von: Plisiecki, Hubert
Veröffentlicht: (2026)
von: Plisiecki, Hubert
Veröffentlicht: (2026)
Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation
von: Carandang, Kristine Ann M., et al.
Veröffentlicht: (2025)
von: Carandang, Kristine Ann M., et al.
Veröffentlicht: (2025)
Towards Explainability in Legal Outcome Prediction Models
von: Valvoda, Josef, et al.
Veröffentlicht: (2024)
von: Valvoda, Josef, et al.
Veröffentlicht: (2024)
Towards a Diagnostic and Predictive Evaluation Methodology for Sequence Labeling Tasks
von: Alvarez-Mellado, Elena, et al.
Veröffentlicht: (2026)
von: Alvarez-Mellado, Elena, et al.
Veröffentlicht: (2026)
The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination
von: Zhang, Yuji, et al.
Veröffentlicht: (2025)
von: Zhang, Yuji, et al.
Veröffentlicht: (2025)
Prediction-powered estimators for finite population statistics in highly imbalanced textual data: Public hate crime estimation
von: Waldetoft, Hannes, et al.
Veröffentlicht: (2025)
von: Waldetoft, Hannes, et al.
Veröffentlicht: (2025)
Predictive Data Selection: The Data That Predicts Is the Data That Teaches
von: Shum, Kashun, et al.
Veröffentlicht: (2025)
von: Shum, Kashun, et al.
Veröffentlicht: (2025)
Predicting Through Generation: Why Generation Is Better for Prediction
von: Kowsher, Md, et al.
Veröffentlicht: (2025)
von: Kowsher, Md, et al.
Veröffentlicht: (2025)
Towards Predictive Communication with Brain-Computer Interfaces integrating Large Language Models
von: Caria, Andrea
Veröffentlicht: (2024)
von: Caria, Andrea
Veröffentlicht: (2024)
Towards Predicting Any Human Trajectory In Context
von: Fujii, Ryo, et al.
Veröffentlicht: (2025)
von: Fujii, Ryo, et al.
Veröffentlicht: (2025)
DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
von: Habba, Eliya, et al.
Veröffentlicht: (2025)
von: Habba, Eliya, et al.
Veröffentlicht: (2025)
Continuous Risk Prediction
von: Dai, Yi
Veröffentlicht: (2024)
von: Dai, Yi
Veröffentlicht: (2024)
Towards Reducing Diagnostic Errors with Interpretable Risk Prediction
von: McInerney, Denis Jered, et al.
Veröffentlicht: (2024)
von: McInerney, Denis Jered, et al.
Veröffentlicht: (2024)
Towards a Psychology of Machines: Large Language Models Predict Human Memory
von: Huff, Markus, et al.
Veröffentlicht: (2024)
von: Huff, Markus, et al.
Veröffentlicht: (2024)
Towards Reliable and Interpretable Traffic Crash Pattern Prediction and Safety Interventions Using Customized Large Language Models
von: Zhao, Yang, et al.
Veröffentlicht: (2025)
von: Zhao, Yang, et al.
Veröffentlicht: (2025)
Towards Robust and Fair Next Visit Diagnosis Prediction under Noisy Clinical Notes with Large Language Models
von: Koo, Heejoon
Veröffentlicht: (2025)
von: Koo, Heejoon
Veröffentlicht: (2025)
Beyond the Next Token: Towards Prompt-Robust Zero-Shot Classification via Efficient Multi-Token Prediction
von: Qian, Junlang, et al.
Veröffentlicht: (2025)
von: Qian, Junlang, et al.
Veröffentlicht: (2025)
Towards Better Graph-based Cross-document Relation Extraction via Non-bridge Entity Enhancement and Prediction Debiasing
von: Yue, Hao, et al.
Veröffentlicht: (2024)
von: Yue, Hao, et al.
Veröffentlicht: (2024)
The Craft of Selective Prediction: Towards Reliable Case Outcome Classification -- An Empirical Study on European Court of Human Rights Cases
von: Santosh, T. Y. S. S., et al.
Veröffentlicht: (2024)
von: Santosh, T. Y. S. S., et al.
Veröffentlicht: (2024)
Towards Auto-Regressive Next-Token Prediction: In-Context Learning Emerges from Generalization
von: Gong, Zixuan, et al.
Veröffentlicht: (2025)
von: Gong, Zixuan, et al.
Veröffentlicht: (2025)
Beyond prompt brittleness: Evaluating the reliability and consistency of political worldviews in LLMs
von: Ceron, Tanise, et al.
Veröffentlicht: (2024)
von: Ceron, Tanise, et al.
Veröffentlicht: (2024)
Legal Fact Prediction: The Missing Piece in Legal Judgment Prediction
von: Liu, Junkai, et al.
Veröffentlicht: (2024)
von: Liu, Junkai, et al.
Veröffentlicht: (2024)
Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining
von: Peng, Bo, et al.
Veröffentlicht: (2026)
von: Peng, Bo, et al.
Veröffentlicht: (2026)
Non-Linear Scoring Model for Translation Quality Evaluation
von: Gladkoff, Serge, et al.
Veröffentlicht: (2025)
von: Gladkoff, Serge, et al.
Veröffentlicht: (2025)
Exploring Fact Memorization and Style Imitation in LLMs Using QLoRA: An Experimental Study and Quality Assessment Methods
von: Vyborov, Eugene, et al.
Veröffentlicht: (2024)
von: Vyborov, Eugene, et al.
Veröffentlicht: (2024)
Ideology Prediction of German Political Texts
von: Schneider, Sinclair, et al.
Veröffentlicht: (2026)
von: Schneider, Sinclair, et al.
Veröffentlicht: (2026)
Reward Prediction with Factorized World States
von: Shen, Yijun, et al.
Veröffentlicht: (2026)
von: Shen, Yijun, et al.
Veröffentlicht: (2026)
Evaluating the Homogeneity of Keyphrase Prediction Models
von: Houbre, Maël, et al.
Veröffentlicht: (2026)
von: Houbre, Maël, et al.
Veröffentlicht: (2026)
Promptly Predicting Structures: The Return of Inference
von: Mehta, Maitrey, et al.
Veröffentlicht: (2024)
von: Mehta, Maitrey, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
How Much Noise Can BERT Handle? Insights from Multilingual Sentence Difficulty Detection
von: Khallaf, Nouran, et al.
Veröffentlicht: (2026) -
Reading Between the Lines: A dataset and a study on why some texts are tougher than others
von: Khallaf, Nouran, et al.
Veröffentlicht: (2025) -
Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification
von: Hilasaca, Kenji, et al.
Veröffentlicht: (2026) -
Can LLM Reasoning Be Trusted? A Comparative Study: Using Human Benchmarking on Statistical Tasks
von: Nagarkar, Crish, et al.
Veröffentlicht: (2026) -
Controlling Out-of-Domain Gaps in LLMs for Genre Classification and Generated Text Detection
von: Roussinov, Dmitri, et al.
Veröffentlicht: (2024)