Calibrating Model-Based Evaluation Metrics for Summarization
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Hongye, Brahma, Dhanajit, Henao, Ricardo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SAGE: A Novelty Gate for Efficient Memory Evolution in Agentic LLMs
by: Wang, Sijia, et al.
Published: (2026)
by: Wang, Sijia, et al.
Published: (2026)
Learning to Substitute Words with Model-based Score Ranking
by: Liu, Hongye, et al.
Published: (2025)
by: Liu, Hongye, et al.
Published: (2025)
Learning to Control Summaries with Score Ranking
by: Liu, Hongye, et al.
Published: (2026)
by: Liu, Hongye, et al.
Published: (2026)
APPLS: Evaluating Evaluation Metrics for Plain Language Summarization
by: Guo, Yue, et al.
Published: (2023)
by: Guo, Yue, et al.
Published: (2023)
CASPR: Automated Evaluation Metric for Contrastive Summarization
by: Ananthamurugan, Nirupan, et al.
Published: (2024)
by: Ananthamurugan, Nirupan, et al.
Published: (2024)
Coupling Generative Modeling and an Autoencoder with the Causal Bridge
by: Meng, Ruolin, et al.
Published: (2025)
by: Meng, Ruolin, et al.
Published: (2025)
Labels have Human Values: Value Calibration of Subjective Tasks
by: Parappan, Mohammed Fayiz, et al.
Published: (2026)
by: Parappan, Mohammed Fayiz, et al.
Published: (2026)
A LongFormer-Based Framework for Accurate and Efficient Medical Text Summarization
by: Sun, Dan, et al.
Published: (2025)
by: Sun, Dan, et al.
Published: (2025)
Polarity Calibration for Opinion Summarization
by: Lei, Yuanyuan, et al.
Published: (2024)
by: Lei, Yuanyuan, et al.
Published: (2024)
Calibration of Large Language Models on Code Summarization
by: Virk, Yuvraj, et al.
Published: (2024)
by: Virk, Yuvraj, et al.
Published: (2024)
Beyond N-Grams: Rethinking Evaluation Metrics and Strategies for Multilingual Abstractive Summarization
by: Mondshine, Itai, et al.
Published: (2025)
by: Mondshine, Itai, et al.
Published: (2025)
Rethinking Scientific Summarization Evaluation: Grounding Explainable Metrics on Facet-aware Benchmark
by: Chen, Xiuying, et al.
Published: (2024)
by: Chen, Xiuying, et al.
Published: (2024)
An Automated Length-Aware Quality Metric for Summarization
by: Foland, Andrew D.
Published: (2025)
by: Foland, Andrew D.
Published: (2025)
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages
by: Yari, Amir Hossein, et al.
Published: (2025)
by: Yari, Amir Hossein, et al.
Published: (2025)
Faithful Model Evaluation for Model-Based Metrics
by: Goyal, Palash, et al.
Published: (2023)
by: Goyal, Palash, et al.
Published: (2023)
PlainQAFact: Retrieval-augmented Factual Consistency Evaluation Metric for Biomedical Plain Language Summarization
by: You, Zhiwen, et al.
Published: (2025)
by: You, Zhiwen, et al.
Published: (2025)
Enhancing Argument Summarization: Prioritizing Exhaustiveness in Key Point Generation and Introducing an Automatic Coverage Evaluation Metric
by: Khosravani, Mohammad, et al.
Published: (2024)
by: Khosravani, Mohammad, et al.
Published: (2024)
Comparative Study of Zero-Shot Cross-Lingual Transfer for Bodo POS and NER Tagging Using Gemini 2.0 Flash Thinking Experimental Model
by: Narzary, Sanjib, et al.
Published: (2025)
by: Narzary, Sanjib, et al.
Published: (2025)
Mitigating the Impact of Reference Quality on Evaluation of Summarization Systems with Reference-Free Metrics
by: Gigant, Théo, et al.
Published: (2024)
by: Gigant, Théo, et al.
Published: (2024)
Redundancy Aware Multi-Reference Based Gainwise Evaluation of Extractive Summarization
by: Akter, Mousumi, et al.
Published: (2023)
by: Akter, Mousumi, et al.
Published: (2023)
Argument Summarization and its Evaluation in the Era of Large Language Models
by: Altemeyer, Moritz, et al.
Published: (2025)
by: Altemeyer, Moritz, et al.
Published: (2025)
Cross-lingual Cross-temporal Summarization: Dataset, Models, Evaluation
by: Zhang, Ran, et al.
Published: (2023)
by: Zhang, Ran, et al.
Published: (2023)
Full-ECE: A Metric For Token-level Calibration on Large Language Models
by: Liu, Han, et al.
Published: (2024)
by: Liu, Han, et al.
Published: (2024)
What's under the hood: Investigating Automatic Metrics on Meeting Summarization
by: Kirstein, Frederic, et al.
Published: (2024)
by: Kirstein, Frederic, et al.
Published: (2024)
Multi-Dimensional Evaluation of Text Summarization with In-Context Learning
by: Jain, Sameer, et al.
Published: (2023)
by: Jain, Sameer, et al.
Published: (2023)
SteerEval: Inference-time Interventions Strengthen Multilingual Generalization in Neural Summarization Metrics
by: Casola, Silvia, et al.
Published: (2026)
by: Casola, Silvia, et al.
Published: (2026)
CCSBench: Evaluating Compositional Controllability in LLMs for Scientific Document Summarization
by: Ding, Yixi, et al.
Published: (2024)
by: Ding, Yixi, et al.
Published: (2024)
On the Role of Summary Content Units in Text Summarization Evaluation
by: Nawrath, Marcel, et al.
Published: (2024)
by: Nawrath, Marcel, et al.
Published: (2024)
On Learning to Summarize with Large Language Models as References
by: Liu, Yixin, et al.
Published: (2023)
by: Liu, Yixin, et al.
Published: (2023)
GRACE: A Granular Benchmark for Evaluating Model Calibration against Human Calibration
by: Sung, Yoo Yeon, et al.
Published: (2025)
by: Sung, Yoo Yeon, et al.
Published: (2025)
CREAM: Comparison-Based Reference-Free ELO-Ranked Automatic Evaluation for Meeting Summarization
by: Gong, Ziwei, et al.
Published: (2024)
by: Gong, Ziwei, et al.
Published: (2024)
Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization
by: Liu, Yixin, et al.
Published: (2023)
by: Liu, Yixin, et al.
Published: (2023)
Evaluating Small Language Models for News Summarization: Implications and Factors Influencing Performance
by: Xu, Borui, et al.
Published: (2025)
by: Xu, Borui, et al.
Published: (2025)
Reading Subtext: Evaluating Large Language Models on Short Story Summarization with Writers
by: Subbiah, Melanie, et al.
Published: (2024)
by: Subbiah, Melanie, et al.
Published: (2024)
AdaptEval: Evaluating Large Language Models on Domain Adaptation for Text Summarization
by: Afzal, Anum, et al.
Published: (2024)
by: Afzal, Anum, et al.
Published: (2024)
Performance Evaluation of Large Language Models in Bangla Consumer Health Query Summarization
by: Abrar, Ajwad, et al.
Published: (2025)
by: Abrar, Ajwad, et al.
Published: (2025)
Evaluating LLMs and Pre-trained Models for Text Summarization Across Diverse Datasets
by: Rehman, Tohida, et al.
Published: (2025)
by: Rehman, Tohida, et al.
Published: (2025)
WPN: An Unlearning Method Based on N-pair Contrastive Learning in Language Models
by: Chen, Guitao, et al.
Published: (2024)
by: Chen, Guitao, et al.
Published: (2024)
Citation-Based Summarization of Landmark Judgments
by: Bindal, Purnima, et al.
Published: (2024)
by: Bindal, Purnima, et al.
Published: (2024)
Modeling Motivated Reasoning in Law: Evaluating Strategic Role Conditioning in LLM Summarization
by: Cho, Eunjung, et al.
Published: (2025)
by: Cho, Eunjung, et al.
Published: (2025)
Similar Items
-
SAGE: A Novelty Gate for Efficient Memory Evolution in Agentic LLMs
by: Wang, Sijia, et al.
Published: (2026) -
Learning to Substitute Words with Model-based Score Ranking
by: Liu, Hongye, et al.
Published: (2025) -
Learning to Control Summaries with Score Ranking
by: Liu, Hongye, et al.
Published: (2026) -
APPLS: Evaluating Evaluation Metrics for Plain Language Summarization
by: Guo, Yue, et al.
Published: (2023) -
CASPR: Automated Evaluation Metric for Contrastive Summarization
by: Ananthamurugan, Nirupan, et al.
Published: (2024)