Can We Trust the Performance Evaluation of Uncertainty Estimation Methods in Text Summarization?
Fuente:
arXiv
Salvato in:
| Autori principali: | He, Jianfeng, Yang, Runing, Yu, Linlin, Li, Changbin, Jia, Ruoxi, Chen, Feng, Jin, Ming, Lu, Chang-Tien |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Uncertainty Estimation on Sequential Labeling via Uncertainty Transmission
di: He, Jianfeng, et al.
Pubblicazione: (2023)
di: He, Jianfeng, et al.
Pubblicazione: (2023)
LLMs Can Plan Only If We Tell Them
di: Sel, Bilgehan, et al.
Pubblicazione: (2025)
di: Sel, Bilgehan, et al.
Pubblicazione: (2025)
InternalInspector $I^2$: Robust Confidence Estimation in LLMs through Internal States
di: Beigi, Mohammad, et al.
Pubblicazione: (2024)
di: Beigi, Mohammad, et al.
Pubblicazione: (2024)
Don't Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech Detection
di: Zhang, Min, et al.
Pubblicazione: (2024)
di: Zhang, Min, et al.
Pubblicazione: (2024)
Can We Trust LLM Detectors?
di: Sandhan, Jivnesh, et al.
Pubblicazione: (2026)
di: Sandhan, Jivnesh, et al.
Pubblicazione: (2026)
LLM-REVal: Can We Trust LLM Reviewers Yet?
di: Li, Rui, et al.
Pubblicazione: (2025)
di: Li, Rui, et al.
Pubblicazione: (2025)
DUAL: Diversity and Uncertainty Active Learning for Text Summarization
di: Giouroukis, Petros Stylianos, et al.
Pubblicazione: (2025)
di: Giouroukis, Petros Stylianos, et al.
Pubblicazione: (2025)
A Comparative Study of Quality Evaluation Methods for Text Summarization
di: Nguyen, Huyen, et al.
Pubblicazione: (2024)
di: Nguyen, Huyen, et al.
Pubblicazione: (2024)
Can We Trust LLMs on Memristors? Diving into Reasoning Ability under Non-Ideality
di: Wu, Taiqiang, et al.
Pubblicazione: (2026)
di: Wu, Taiqiang, et al.
Pubblicazione: (2026)
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?
di: Wang, Leyao, et al.
Pubblicazione: (2026)
di: Wang, Leyao, et al.
Pubblicazione: (2026)
TCMD: A Traditional Chinese Medicine QA Dataset for Evaluating Large Language Models
di: Yu, Ping, et al.
Pubblicazione: (2024)
di: Yu, Ping, et al.
Pubblicazione: (2024)
Exploring the Deceptive Power of LLM-Generated Fake News: A Study of Real-World Detection Challenges
di: Sun, Yanshen, et al.
Pubblicazione: (2024)
di: Sun, Yanshen, et al.
Pubblicazione: (2024)
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
di: Xiong, Miao, et al.
Pubblicazione: (2023)
di: Xiong, Miao, et al.
Pubblicazione: (2023)
Comprehensive Manuscript Assessment with Text Summarization Using 69707 articles
di: Sun, Qichen, et al.
Pubblicazione: (2025)
di: Sun, Qichen, et al.
Pubblicazione: (2025)
A Systematic Survey of Text Summarization: From Statistical Methods to Large Language Models
di: Zhang, Haopeng, et al.
Pubblicazione: (2024)
di: Zhang, Haopeng, et al.
Pubblicazione: (2024)
Multi-Dimensional Evaluation of Text Summarization with In-Context Learning
di: Jain, Sameer, et al.
Pubblicazione: (2023)
di: Jain, Sameer, et al.
Pubblicazione: (2023)
On the Role of Summary Content Units in Text Summarization Evaluation
di: Nawrath, Marcel, et al.
Pubblicazione: (2024)
di: Nawrath, Marcel, et al.
Pubblicazione: (2024)
Can We Trust LLMs? Mitigate Overconfidence Bias in LLMs through Knowledge Transfer
di: Yang, Haoyan, et al.
Pubblicazione: (2024)
di: Yang, Haoyan, et al.
Pubblicazione: (2024)
MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization
di: Liu, Yinhong, et al.
Pubblicazione: (2025)
di: Liu, Yinhong, et al.
Pubblicazione: (2025)
When Can We Trust LLM Graders? Calibrating Confidence for Automated Assessment
di: Ferrer, Robinson, et al.
Pubblicazione: (2026)
di: Ferrer, Robinson, et al.
Pubblicazione: (2026)
Enriching and Controlling Global Semantics for Text Summarization
di: Nguyen, Thong, et al.
Pubblicazione: (2021)
di: Nguyen, Thong, et al.
Pubblicazione: (2021)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
di: Badawi, Abeer, et al.
Pubblicazione: (2025)
di: Badawi, Abeer, et al.
Pubblicazione: (2025)
Survey of Query-based Text Summarization
di: Yu, Hang, et al.
Pubblicazione: (2022)
di: Yu, Hang, et al.
Pubblicazione: (2022)
Rethinking LLM Evaluation: Can We Evaluate LLMs with 200x Less Data?
di: Wang, Shaobo, et al.
Pubblicazione: (2025)
di: Wang, Shaobo, et al.
Pubblicazione: (2025)
Scientific Opinion Summarization: Paper Meta-review Generation Dataset, Methods, and Evaluation
di: Zeng, Qi, et al.
Pubblicazione: (2023)
di: Zeng, Qi, et al.
Pubblicazione: (2023)
Can We Trust LLMs for Mental Health Screening? Consistency, ASR Robustness, and Evidence Faithfulness
di: Loweimi, Erfan, et al.
Pubblicazione: (2026)
di: Loweimi, Erfan, et al.
Pubblicazione: (2026)
Adapted Large Language Models Can Outperform Medical Experts in Clinical Text Summarization
di: Van Veen, Dave, et al.
Pubblicazione: (2023)
di: Van Veen, Dave, et al.
Pubblicazione: (2023)
End-to-End Chatbot Evaluation with Adaptive Reasoning and Uncertainty Filtering
di: Dang, Nhi, et al.
Pubblicazione: (2026)
di: Dang, Nhi, et al.
Pubblicazione: (2026)
Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
di: Prandi, Matteo, et al.
Pubblicazione: (2025)
di: Prandi, Matteo, et al.
Pubblicazione: (2025)
Text Style Transfer with Parameter-efficient LLM Finetuning and Round-trip Translation
di: Liu, Ruoxi, et al.
Pubblicazione: (2026)
di: Liu, Ruoxi, et al.
Pubblicazione: (2026)
Can We Predict Performance of Large Models across Vision-Language Tasks?
di: Zhao, Qinyu, et al.
Pubblicazione: (2024)
di: Zhao, Qinyu, et al.
Pubblicazione: (2024)
Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models
di: Hobelsberger, Christian, et al.
Pubblicazione: (2025)
di: Hobelsberger, Christian, et al.
Pubblicazione: (2025)
Evaluating Small Language Models for News Summarization: Implications and Factors Influencing Performance
di: Xu, Borui, et al.
Pubblicazione: (2025)
di: Xu, Borui, et al.
Pubblicazione: (2025)
Multi-LLM Text Summarization
di: Fang, Jiangnan, et al.
Pubblicazione: (2024)
di: Fang, Jiangnan, et al.
Pubblicazione: (2024)
Topic-Controllable Summarization: Topic-Aware Evaluation and Transformer Methods
di: Passali, Tatiana, et al.
Pubblicazione: (2022)
di: Passali, Tatiana, et al.
Pubblicazione: (2022)
QAPyramid: Fine-grained Evaluation of Content Selection for Text Summarization
di: Zhang, Shiyue, et al.
Pubblicazione: (2024)
di: Zhang, Shiyue, et al.
Pubblicazione: (2024)
APPLS: Evaluating Evaluation Metrics for Plain Language Summarization
di: Guo, Yue, et al.
Pubblicazione: (2023)
di: Guo, Yue, et al.
Pubblicazione: (2023)
Can LLMs Infer Personality from Real World Conversations?
di: Zhu, Jianfeng, et al.
Pubblicazione: (2025)
di: Zhu, Jianfeng, et al.
Pubblicazione: (2025)
Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models
di: Sel, Bilgehan, et al.
Pubblicazione: (2023)
di: Sel, Bilgehan, et al.
Pubblicazione: (2023)
Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents
di: Al-Tawaha, Ahmad, et al.
Pubblicazione: (2026)
di: Al-Tawaha, Ahmad, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Uncertainty Estimation on Sequential Labeling via Uncertainty Transmission
di: He, Jianfeng, et al.
Pubblicazione: (2023) -
LLMs Can Plan Only If We Tell Them
di: Sel, Bilgehan, et al.
Pubblicazione: (2025) -
InternalInspector $I^2$: Robust Confidence Estimation in LLMs through Internal States
di: Beigi, Mohammad, et al.
Pubblicazione: (2024) -
Don't Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech Detection
di: Zhang, Min, et al.
Pubblicazione: (2024) -
Can We Trust LLM Detectors?
di: Sandhan, Jivnesh, et al.
Pubblicazione: (2026)