Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Dong, Rastogi, Eti, Zhao, Fen, Goyal, Sagar, Naik, Gautam, Rajagopal, Sree Prasanna |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Continued Pretrained LLM Approach for Automatic Medical Note Generation
by: Yuan, Dong, et al.
Published: (2024)
by: Yuan, Dong, et al.
Published: (2024)
Extrinsically-Focused Evaluation of Omissions in Medical Summarization
by: Schumacher, Elliot, et al.
Published: (2023)
by: Schumacher, Elliot, et al.
Published: (2023)
FineSurE: Fine-grained Summarization Evaluation using LLMs
by: Song, Hwanjun, et al.
Published: (2024)
by: Song, Hwanjun, et al.
Published: (2024)
QAPyramid: Fine-grained Evaluation of Content Selection for Text Summarization
by: Zhang, Shiyue, et al.
Published: (2024)
by: Zhang, Shiyue, et al.
Published: (2024)
Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages
by: Min, Hyangsuk, et al.
Published: (2025)
by: Min, Hyangsuk, et al.
Published: (2025)
UniSumEval: Towards Unified, Fine-Grained, Multi-Dimensional Summarization Evaluation for LLMs
by: Lee, Yuho, et al.
Published: (2024)
by: Lee, Yuho, et al.
Published: (2024)
STORYSUMM: Evaluating Faithfulness in Story Summarization
by: Subbiah, Melanie, et al.
Published: (2024)
by: Subbiah, Melanie, et al.
Published: (2024)
Abstractive Summarization of Low resourced Nepali language using Multilingual Transformers
by: Dhakal, Prakash, et al.
Published: (2024)
by: Dhakal, Prakash, et al.
Published: (2024)
ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization
by: Bae, Suyoung, et al.
Published: (2026)
by: Bae, Suyoung, et al.
Published: (2026)
LMUnit: Fine-grained Evaluation with Natural Language Unit Tests
by: Saad-Falcon, Jon, et al.
Published: (2024)
by: Saad-Falcon, Jon, et al.
Published: (2024)
Auto-ARGUE: LLM-Based Report Generation Evaluation
by: Walden, William, et al.
Published: (2025)
by: Walden, William, et al.
Published: (2025)
ByteScience: Bridging Unstructured Scientific Literature and Structured Data with Auto Fine-tuned Large Language Model in Token Granularity
by: Xie, Tong, et al.
Published: (2024)
by: Xie, Tong, et al.
Published: (2024)
AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data
by: Wu, JiaRu, et al.
Published: (2025)
by: Wu, JiaRu, et al.
Published: (2025)
SciQAG: A Framework for Auto-Generated Science Question Answering Dataset with Fine-grained Evaluation
by: Wan, Yuwei, et al.
Published: (2024)
by: Wan, Yuwei, et al.
Published: (2024)
TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization
by: Tang, Liyan, et al.
Published: (2024)
by: Tang, Liyan, et al.
Published: (2024)
$\texttt{COSMIC}$: Mutual Information for Task-Agnostic Summarization Evaluation
by: Darrin, Maxime, et al.
Published: (2024)
by: Darrin, Maxime, et al.
Published: (2024)
A Comparative Study of Quality Evaluation Methods for Text Summarization
by: Nguyen, Huyen, et al.
Published: (2024)
by: Nguyen, Huyen, et al.
Published: (2024)
MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization
by: Liu, Yinhong, et al.
Published: (2025)
by: Liu, Yinhong, et al.
Published: (2025)
Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization
by: Jin, Keyan, et al.
Published: (2025)
by: Jin, Keyan, et al.
Published: (2025)
Policies and Evaluation for Online Meeting Summarization
by: Schneider, Felix, et al.
Published: (2025)
by: Schneider, Felix, et al.
Published: (2025)
Evaluating Large Language Models on Financial Report Summarization: An Empirical Study
by: Yang, Xinqi, et al.
Published: (2024)
by: Yang, Xinqi, et al.
Published: (2024)
Discourse-Driven Evaluation: Unveiling Factual Inconsistency in Long Document Summarization
by: Zhong, Yang, et al.
Published: (2025)
by: Zhong, Yang, et al.
Published: (2025)
PersonaMatrix: A Recipe for Persona-Aware Evaluation of Legal Summarization
by: Pang, Tsz Fung, et al.
Published: (2025)
by: Pang, Tsz Fung, et al.
Published: (2025)
SEval-Ex: A Statement-Level Framework for Explainable Summarization Evaluation
by: Herserant, Tanguy, et al.
Published: (2025)
by: Herserant, Tanguy, et al.
Published: (2025)
Prompt Sentiment: The Catalyst for LLM Change
by: Gandhi, Vishal, et al.
Published: (2025)
by: Gandhi, Vishal, et al.
Published: (2025)
Leveraging the Power of LLMs: A Fine-Tuning Approach for High-Quality Aspect-Based Summarization
by: Mullick, Ankan, et al.
Published: (2024)
by: Mullick, Ankan, et al.
Published: (2024)
Evaluating Small Language Models for News Summarization: Implications and Factors Influencing Performance
by: Xu, Borui, et al.
Published: (2025)
by: Xu, Borui, et al.
Published: (2025)
AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
by: Ryan, Michael J., et al.
Published: (2025)
by: Ryan, Michael J., et al.
Published: (2025)
Fine-Tuned Language Models for Domain-Specific Summarization and Tagging
by: Wang, Jun, et al.
Published: (2025)
by: Wang, Jun, et al.
Published: (2025)
Learning to Summarize from LLM-generated Feedback
by: Song, Hwanjun, et al.
Published: (2024)
by: Song, Hwanjun, et al.
Published: (2024)
TALEC: Teach Your LLM to Evaluate in Specific Domain with In-house Criteria by Criteria Division and Zero-shot Plus Few-shot
by: Zhang, Kaiqi, et al.
Published: (2024)
by: Zhang, Kaiqi, et al.
Published: (2024)
Evaluation of Large Language Models for Summarization Tasks in the Medical Domain: A Narrative Review
by: Croxford, Emma, et al.
Published: (2024)
by: Croxford, Emma, et al.
Published: (2024)
An Evaluation of Large Language Models on Text Summarization Tasks Using Prompt Engineering Techniques
by: Aly, Walid Mohamed, et al.
Published: (2025)
by: Aly, Walid Mohamed, et al.
Published: (2025)
CurLL: A Developmental Framework to Evaluate Continual Learning in Language Models
by: Kalyan, Pavan, et al.
Published: (2025)
by: Kalyan, Pavan, et al.
Published: (2025)
LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning
by: Huang, Wei, et al.
Published: (2026)
by: Huang, Wei, et al.
Published: (2026)
AutoMix: Automatically Mixing Language Models
by: Aggarwal, Pranjal, et al.
Published: (2023)
by: Aggarwal, Pranjal, et al.
Published: (2023)
Towards Coarse-to-Fine Evaluation of Inference Efficiency for Large Language Models
by: Chen, Yushuo, et al.
Published: (2024)
by: Chen, Yushuo, et al.
Published: (2024)
ReFeR: Improving Evaluation and Reasoning through Hierarchy of Models
by: Narsupalli, Yaswanth, et al.
Published: (2024)
by: Narsupalli, Yaswanth, et al.
Published: (2024)
Prompt Recursive Search: A Living Framework with Adaptive Growth in LLM Auto-Prompting
by: Zhao, Xiangyu, et al.
Published: (2024)
by: Zhao, Xiangyu, et al.
Published: (2024)
On the Benefits of Fine-Grained Loss Truncation: A Case Study on Factuality in Summarization
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2024)
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2024)
Similar Items
-
A Continued Pretrained LLM Approach for Automatic Medical Note Generation
by: Yuan, Dong, et al.
Published: (2024) -
Extrinsically-Focused Evaluation of Omissions in Medical Summarization
by: Schumacher, Elliot, et al.
Published: (2023) -
FineSurE: Fine-grained Summarization Evaluation using LLMs
by: Song, Hwanjun, et al.
Published: (2024) -
QAPyramid: Fine-grained Evaluation of Content Selection for Text Summarization
by: Zhang, Shiyue, et al.
Published: (2024) -
Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages
by: Min, Hyangsuk, et al.
Published: (2025)