Are Generative Models Underconfident? Better Quality Estimation with Boosted Model Probability
Fuente:
arXiv
Salvato in:
| Autori principali: | Dinh, Tu Anh, Niehues, Jan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Quality Estimation with $k$-nearest Neighbors and Automatic Evaluation for Model-specific Quality Estimation
di: Dinh, Tu Anh, et al.
Pubblicazione: (2024)
di: Dinh, Tu Anh, et al.
Pubblicazione: (2024)
Sigmoid Head for Quality Estimation under Language Ambiguity
di: Dinh, Tu Anh, et al.
Pubblicazione: (2026)
di: Dinh, Tu Anh, et al.
Pubblicazione: (2026)
Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons
di: Sandan, Isik Baran, et al.
Pubblicazione: (2025)
di: Sandan, Isik Baran, et al.
Pubblicazione: (2025)
COMET-poly: Machine Translation Metric Grounded in Other Candidates
di: Züfle, Maike, et al.
Pubblicazione: (2025)
di: Züfle, Maike, et al.
Pubblicazione: (2025)
SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
di: Dinh, Tu Anh, et al.
Pubblicazione: (2024)
di: Dinh, Tu Anh, et al.
Pubblicazione: (2024)
Tokenization and Morphology in Multilingual Language Models: A Comparative Analysis of mT5 and ByT5
di: Dang, Thao Anh, et al.
Pubblicazione: (2024)
di: Dang, Thao Anh, et al.
Pubblicazione: (2024)
Surprise Calibration for Better In-Context Learning
di: Tan, Zhihang, et al.
Pubblicazione: (2025)
di: Tan, Zhihang, et al.
Pubblicazione: (2025)
Is Less More? Quality, Quantity and Context in Idiom Processing with Natural Language Models
di: Knietaite, Agne, et al.
Pubblicazione: (2024)
di: Knietaite, Agne, et al.
Pubblicazione: (2024)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
di: Collado-Montañez, Jaime, et al.
Pubblicazione: (2025)
di: Collado-Montañez, Jaime, et al.
Pubblicazione: (2025)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
di: Peters, Sydney, et al.
Pubblicazione: (2025)
di: Peters, Sydney, et al.
Pubblicazione: (2025)
Towards Human Understanding of Paraphrase Types in Large Language Models
di: Meier, Dominik, et al.
Pubblicazione: (2024)
di: Meier, Dominik, et al.
Pubblicazione: (2024)
Lightweight Connective Detection Using Gradient Boosting
di: Er, Mustafa Erolcan, et al.
Pubblicazione: (2024)
di: Er, Mustafa Erolcan, et al.
Pubblicazione: (2024)
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
di: Stacey, Joe, et al.
Pubblicazione: (2026)
di: Stacey, Joe, et al.
Pubblicazione: (2026)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
What Language Models Know But Don't Say: Non-Generative Prior Extraction for Generalization
di: Rezaeimanesh, Sara, et al.
Pubblicazione: (2026)
di: Rezaeimanesh, Sara, et al.
Pubblicazione: (2026)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
di: Tong, Jingqi, et al.
Pubblicazione: (2025)
di: Tong, Jingqi, et al.
Pubblicazione: (2025)
Encoder-Decoder Framework for Interactive Free Verses with Generation with Controllable High-Quality Rhyming
di: Pasini, Tommaso, et al.
Pubblicazione: (2024)
di: Pasini, Tommaso, et al.
Pubblicazione: (2024)
German Text Simplification: Finetuning Large Language Models with Semi-Synthetic Data
di: Klöser, Lars, et al.
Pubblicazione: (2024)
di: Klöser, Lars, et al.
Pubblicazione: (2024)
After Retrieval, Before Generation: Enhancing the Trustworthiness of Large Language Models in Retrieval-Augmented Generation
di: Dai, Xinbang, et al.
Pubblicazione: (2025)
di: Dai, Xinbang, et al.
Pubblicazione: (2025)
KinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generation
di: Nzeyimana, Antoine, et al.
Pubblicazione: (2025)
di: Nzeyimana, Antoine, et al.
Pubblicazione: (2025)
Refining Packing and Shuffling Strategies for Enhanced Performance in Generative Language Models
di: Chen, Yanbing, et al.
Pubblicazione: (2024)
di: Chen, Yanbing, et al.
Pubblicazione: (2024)
Textual Entailment is not a Better Bias Metric than Token Probability
di: Felkner, Virginia K., et al.
Pubblicazione: (2025)
di: Felkner, Virginia K., et al.
Pubblicazione: (2025)
Steering Language Models in Multi-Token Generation: A Case Study on Tense and Aspect
di: Klerings, Alina, et al.
Pubblicazione: (2025)
di: Klerings, Alina, et al.
Pubblicazione: (2025)
OPOR-Bench: Evaluating Large Language Models on Online Public Opinion Report Generation
di: Yu, Jinzheng, et al.
Pubblicazione: (2025)
di: Yu, Jinzheng, et al.
Pubblicazione: (2025)
Better and Worse with Scale: How Contextual Entrainment Diverges with Model Size
di: Kukreja, Dikshant, et al.
Pubblicazione: (2026)
di: Kukreja, Dikshant, et al.
Pubblicazione: (2026)
UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking
di: Ahmad, Sarfraz, et al.
Pubblicazione: (2025)
di: Ahmad, Sarfraz, et al.
Pubblicazione: (2025)
Boosting Accuracy and Interpretability in Multilingual Hate Speech Detection Through Layer Freezing and Explainable AI
di: Bilehsavar, Meysam Shirdel, et al.
Pubblicazione: (2026)
di: Bilehsavar, Meysam Shirdel, et al.
Pubblicazione: (2026)
SynDocDis: A Metadata-Driven Framework for Generating Synthetic Physician Discussions Using Large Language Models
di: Rubinstein, Beny, et al.
Pubblicazione: (2026)
di: Rubinstein, Beny, et al.
Pubblicazione: (2026)
Multi-Hierarchical Feature Detection for Large Language Model Generated Text
di: Zhang, Luyan, et al.
Pubblicazione: (2025)
di: Zhang, Luyan, et al.
Pubblicazione: (2025)
Adapting Multilingual Models to Code-Mixed Tasks via Model Merging
di: Kodali, Prashant, et al.
Pubblicazione: (2025)
di: Kodali, Prashant, et al.
Pubblicazione: (2025)
Enhancing Emotion Prediction in News Headlines: Insights from ChatGPT and Seq2Seq Models for Free-Text Generation
di: Gao, Ge, et al.
Pubblicazione: (2024)
di: Gao, Ge, et al.
Pubblicazione: (2024)
Large Language Models Can Better Understand Knowledge Graphs Than We Thought
di: Dai, Xinbang, et al.
Pubblicazione: (2024)
di: Dai, Xinbang, et al.
Pubblicazione: (2024)
Exploring the Maze of Multilingual Modeling
di: Nezhad, Sina Bagheri, et al.
Pubblicazione: (2023)
di: Nezhad, Sina Bagheri, et al.
Pubblicazione: (2023)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
di: Liu, Zhongxin, et al.
Pubblicazione: (2025)
di: Liu, Zhongxin, et al.
Pubblicazione: (2025)
LinkNER: Linking Local Named Entity Recognition Models to Large Language Models using Uncertainty
di: Zhang, Zhen, et al.
Pubblicazione: (2024)
di: Zhang, Zhen, et al.
Pubblicazione: (2024)
BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation
di: The Omnilingual MT Team, et al.
Pubblicazione: (2025)
di: The Omnilingual MT Team, et al.
Pubblicazione: (2025)
Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text
di: Kugler, Kai
Pubblicazione: (2025)
di: Kugler, Kai
Pubblicazione: (2025)
Cross-lingual Human-Preference Alignment for Neural Machine Translation with Direct Quality Optimization
di: Uhlig, Kaden, et al.
Pubblicazione: (2024)
di: Uhlig, Kaden, et al.
Pubblicazione: (2024)
Precise Length Control in Large Language Models
di: Butcher, Bradley, et al.
Pubblicazione: (2024)
di: Butcher, Bradley, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Quality Estimation with $k$-nearest Neighbors and Automatic Evaluation for Model-specific Quality Estimation
di: Dinh, Tu Anh, et al.
Pubblicazione: (2024) -
Sigmoid Head for Quality Estimation under Language Ambiguity
di: Dinh, Tu Anh, et al.
Pubblicazione: (2026) -
Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons
di: Sandan, Isik Baran, et al.
Pubblicazione: (2025) -
COMET-poly: Machine Translation Metric Grounded in Other Candidates
di: Züfle, Maike, et al.
Pubblicazione: (2025) -
SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
di: Dinh, Tu Anh, et al.
Pubblicazione: (2024)