Gespeichert in:
| Hauptverfasser: | He, Jianfeng, Yang, Runing, Yu, Linlin, Li, Changbin, Jia, Ruoxi, Chen, Feng, Jin, Ming, Lu, Chang-Tien |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2406.17274 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Uncertainty Estimation on Sequential Labeling via Uncertainty Transmission
von: He, Jianfeng, et al.
Veröffentlicht: (2023)
von: He, Jianfeng, et al.
Veröffentlicht: (2023)
InternalInspector $I^2$: Robust Confidence Estimation in LLMs through Internal States
von: Beigi, Mohammad, et al.
Veröffentlicht: (2024)
von: Beigi, Mohammad, et al.
Veröffentlicht: (2024)
LLMs Can Plan Only If We Tell Them
von: Sel, Bilgehan, et al.
Veröffentlicht: (2025)
von: Sel, Bilgehan, et al.
Veröffentlicht: (2025)
Don't Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech Detection
von: Zhang, Min, et al.
Veröffentlicht: (2024)
von: Zhang, Min, et al.
Veröffentlicht: (2024)
Can We Trust LLM Detectors?
von: Sandhan, Jivnesh, et al.
Veröffentlicht: (2026)
von: Sandhan, Jivnesh, et al.
Veröffentlicht: (2026)
LLM-REVal: Can We Trust LLM Reviewers Yet?
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
A Comparative Study of Quality Evaluation Methods for Text Summarization
von: Nguyen, Huyen, et al.
Veröffentlicht: (2024)
von: Nguyen, Huyen, et al.
Veröffentlicht: (2024)
Exploring the Deceptive Power of LLM-Generated Fake News: A Study of Real-World Detection Challenges
von: Sun, Yanshen, et al.
Veröffentlicht: (2024)
von: Sun, Yanshen, et al.
Veröffentlicht: (2024)
DUAL: Diversity and Uncertainty Active Learning for Text Summarization
von: Giouroukis, Petros Stylianos, et al.
Veröffentlicht: (2025)
von: Giouroukis, Petros Stylianos, et al.
Veröffentlicht: (2025)
TCMD: A Traditional Chinese Medicine QA Dataset for Evaluating Large Language Models
von: Yu, Ping, et al.
Veröffentlicht: (2024)
von: Yu, Ping, et al.
Veröffentlicht: (2024)
Can We Trust LLMs on Memristors? Diving into Reasoning Ability under Non-Ideality
von: Wu, Taiqiang, et al.
Veröffentlicht: (2026)
von: Wu, Taiqiang, et al.
Veröffentlicht: (2026)
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?
von: Wang, Leyao, et al.
Veröffentlicht: (2026)
von: Wang, Leyao, et al.
Veröffentlicht: (2026)
Comprehensive Manuscript Assessment with Text Summarization Using 69707 articles
von: Sun, Qichen, et al.
Veröffentlicht: (2025)
von: Sun, Qichen, et al.
Veröffentlicht: (2025)
MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization
von: Liu, Yinhong, et al.
Veröffentlicht: (2025)
von: Liu, Yinhong, et al.
Veröffentlicht: (2025)
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
von: Xiong, Miao, et al.
Veröffentlicht: (2023)
von: Xiong, Miao, et al.
Veröffentlicht: (2023)
A Systematic Survey of Text Summarization: From Statistical Methods to Large Language Models
von: Zhang, Haopeng, et al.
Veröffentlicht: (2024)
von: Zhang, Haopeng, et al.
Veröffentlicht: (2024)
When Can We Trust LLM Graders? Calibrating Confidence for Automated Assessment
von: Ferrer, Robinson, et al.
Veröffentlicht: (2026)
von: Ferrer, Robinson, et al.
Veröffentlicht: (2026)
Can We Trust LLMs? Mitigate Overconfidence Bias in LLMs through Knowledge Transfer
von: Yang, Haoyan, et al.
Veröffentlicht: (2024)
von: Yang, Haoyan, et al.
Veröffentlicht: (2024)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
von: Badawi, Abeer, et al.
Veröffentlicht: (2025)
von: Badawi, Abeer, et al.
Veröffentlicht: (2025)
Multi-Dimensional Evaluation of Text Summarization with In-Context Learning
von: Jain, Sameer, et al.
Veröffentlicht: (2023)
von: Jain, Sameer, et al.
Veröffentlicht: (2023)
On the Role of Summary Content Units in Text Summarization Evaluation
von: Nawrath, Marcel, et al.
Veröffentlicht: (2024)
von: Nawrath, Marcel, et al.
Veröffentlicht: (2024)
Rethinking LLM Evaluation: Can We Evaluate LLMs with 200x Less Data?
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
Enriching and Controlling Global Semantics for Text Summarization
von: Nguyen, Thong, et al.
Veröffentlicht: (2021)
von: Nguyen, Thong, et al.
Veröffentlicht: (2021)
Survey of Query-based Text Summarization
von: Yu, Hang, et al.
Veröffentlicht: (2022)
von: Yu, Hang, et al.
Veröffentlicht: (2022)
End-to-End Chatbot Evaluation with Adaptive Reasoning and Uncertainty Filtering
von: Dang, Nhi, et al.
Veröffentlicht: (2026)
von: Dang, Nhi, et al.
Veröffentlicht: (2026)
Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models
von: Sel, Bilgehan, et al.
Veröffentlicht: (2023)
von: Sel, Bilgehan, et al.
Veröffentlicht: (2023)
Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents
von: Al-Tawaha, Ahmad, et al.
Veröffentlicht: (2026)
von: Al-Tawaha, Ahmad, et al.
Veröffentlicht: (2026)
Text Style Transfer with Parameter-efficient LLM Finetuning and Round-trip Translation
von: Liu, Ruoxi, et al.
Veröffentlicht: (2026)
von: Liu, Ruoxi, et al.
Veröffentlicht: (2026)
Scientific Opinion Summarization: Paper Meta-review Generation Dataset, Methods, and Evaluation
von: Zeng, Qi, et al.
Veröffentlicht: (2023)
von: Zeng, Qi, et al.
Veröffentlicht: (2023)
Can We Trust LLMs for Mental Health Screening? Consistency, ASR Robustness, and Evidence Faithfulness
von: Loweimi, Erfan, et al.
Veröffentlicht: (2026)
von: Loweimi, Erfan, et al.
Veröffentlicht: (2026)
Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
von: Prandi, Matteo, et al.
Veröffentlicht: (2025)
von: Prandi, Matteo, et al.
Veröffentlicht: (2025)
Can We Predict Performance of Large Models across Vision-Language Tasks?
von: Zhao, Qinyu, et al.
Veröffentlicht: (2024)
von: Zhao, Qinyu, et al.
Veröffentlicht: (2024)
Adapted Large Language Models Can Outperform Medical Experts in Clinical Text Summarization
von: Van Veen, Dave, et al.
Veröffentlicht: (2023)
von: Van Veen, Dave, et al.
Veröffentlicht: (2023)
Evaluating Small Language Models for News Summarization: Implications and Factors Influencing Performance
von: Xu, Borui, et al.
Veröffentlicht: (2025)
von: Xu, Borui, et al.
Veröffentlicht: (2025)
Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models
von: Hobelsberger, Christian, et al.
Veröffentlicht: (2025)
von: Hobelsberger, Christian, et al.
Veröffentlicht: (2025)
Can LLMs Infer Personality from Real World Conversations?
von: Zhu, Jianfeng, et al.
Veröffentlicht: (2025)
von: Zhu, Jianfeng, et al.
Veröffentlicht: (2025)
Multi-LLM Text Summarization
von: Fang, Jiangnan, et al.
Veröffentlicht: (2024)
von: Fang, Jiangnan, et al.
Veröffentlicht: (2024)
QAPyramid: Fine-grained Evaluation of Content Selection for Text Summarization
von: Zhang, Shiyue, et al.
Veröffentlicht: (2024)
von: Zhang, Shiyue, et al.
Veröffentlicht: (2024)
APPLS: Evaluating Evaluation Metrics for Plain Language Summarization
von: Guo, Yue, et al.
Veröffentlicht: (2023)
von: Guo, Yue, et al.
Veröffentlicht: (2023)
Topic-Controllable Summarization: Topic-Aware Evaluation and Transformer Methods
von: Passali, Tatiana, et al.
Veröffentlicht: (2022)
von: Passali, Tatiana, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
Uncertainty Estimation on Sequential Labeling via Uncertainty Transmission
von: He, Jianfeng, et al.
Veröffentlicht: (2023) -
InternalInspector $I^2$: Robust Confidence Estimation in LLMs through Internal States
von: Beigi, Mohammad, et al.
Veröffentlicht: (2024) -
LLMs Can Plan Only If We Tell Them
von: Sel, Bilgehan, et al.
Veröffentlicht: (2025) -
Don't Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech Detection
von: Zhang, Min, et al.
Veröffentlicht: (2024) -
Can We Trust LLM Detectors?
von: Sandhan, Jivnesh, et al.
Veröffentlicht: (2026)