Real-World Summarization: When Evaluation Reaches Its Limits
Fuente:
arXiv
Saved in:
| Main Authors: | Schmidtová, Patrícia, Dušek, Ondřej, Mahamood, Saad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Automatic Metrics in Natural Language Generation: A Survey of Current Evaluation Practices
by: Schmidtová, Patrícia, et al.
Published: (2024)
by: Schmidtová, Patrícia, et al.
Published: (2024)
factgenie: A Framework for Span-based Evaluation of Generated Texts
by: Kasner, Zdeněk, et al.
Published: (2024)
by: Kasner, Zdeněk, et al.
Published: (2024)
Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs
by: Balloccu, Simone, et al.
Published: (2024)
by: Balloccu, Simone, et al.
Published: (2024)
LLMs as Span Annotators: A Comparative Study of LLMs and Humans
by: Kasner, Zdeněk, et al.
Published: (2025)
by: Kasner, Zdeněk, et al.
Published: (2025)
Sentence Embeddings as an intermediate target in end-to-end summarisation
by: Zembrzuski, Maciej, et al.
Published: (2025)
by: Zembrzuski, Maciej, et al.
Published: (2025)
FreshTab: Sourcing Fresh Data for Table-to-Text Generation Evaluation
by: Onderková, Kristýna, et al.
Published: (2025)
by: Onderková, Kristýna, et al.
Published: (2025)
OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs
by: Kartáč, Ivan, et al.
Published: (2025)
by: Kartáč, Ivan, et al.
Published: (2025)
AnimatedLLM: Explaining LLMs with Interactive Visualizations
by: Kasner, Zdeněk, et al.
Published: (2025)
by: Kasner, Zdeněk, et al.
Published: (2025)
Text Style Transfer: An Introductory Overview
by: Mukherjee, Sourabrata, et al.
Published: (2024)
by: Mukherjee, Sourabrata, et al.
Published: (2024)
LEEETs-Dial: Linguistic Entrainment in End-to-End Task-oriented Dialogue systems
by: Kumar, Nalin, et al.
Published: (2023)
by: Kumar, Nalin, et al.
Published: (2023)
Beyond Traditional Benchmarks: Analyzing Behaviors of Open LLMs on Data-to-Text Generation
by: Kasner, Zdeněk, et al.
Published: (2024)
by: Kasner, Zdeněk, et al.
Published: (2024)
LLM Agents Implement an NLG System from Scratch: Building Interpretable Rule-Based RDF-to-Text Generators
by: Lango, Mateusz, et al.
Published: (2025)
by: Lango, Mateusz, et al.
Published: (2025)
When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition
by: Li, Karen Jia-Hui, et al.
Published: (2025)
by: Li, Karen Jia-Hui, et al.
Published: (2025)
SRS-Stories: Vocabulary-constrained multilingual story generation for language learning
by: Kamzela, Wiktor, et al.
Published: (2025)
by: Kamzela, Wiktor, et al.
Published: (2025)
Strategies for Span Labeling with Large Language Models
by: Semin, Danil, et al.
Published: (2026)
by: Semin, Danil, et al.
Published: (2026)
Reasoning Gets Harder for LLMs Inside A Dialogue
by: Kartáč, Ivan, et al.
Published: (2026)
by: Kartáč, Ivan, et al.
Published: (2026)
On the Role of Summary Content Units in Text Summarization Evaluation
by: Nawrath, Marcel, et al.
Published: (2024)
by: Nawrath, Marcel, et al.
Published: (2024)
Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics?
by: Mukherjee, Sourabrata, et al.
Published: (2025)
by: Mukherjee, Sourabrata, et al.
Published: (2025)
Are Large Language Models Actually Good at Text Style Transfer?
by: Mukherjee, Sourabrata, et al.
Published: (2024)
by: Mukherjee, Sourabrata, et al.
Published: (2024)
Leveraging Large Language Models for Building Interpretable Rule-Based Data-to-Text Systems
by: Warczyński, Jędrzej, et al.
Published: (2025)
by: Warczyński, Jędrzej, et al.
Published: (2025)
Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach
by: Wojciechowski, Adam, et al.
Published: (2024)
by: Wojciechowski, Adam, et al.
Published: (2024)
How Do People Quantify Naturally: Evidence from Mandarin Picture Description
by: Zhang, Yayun, et al.
Published: (2026)
by: Zhang, Yayun, et al.
Published: (2026)
A Survey of Text Style Transfer: Applications and Ethical Implications
by: Mukherjee, Sourabrata, et al.
Published: (2024)
by: Mukherjee, Sourabrata, et al.
Published: (2024)
ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization
by: Bae, Suyoung, et al.
Published: (2026)
by: Bae, Suyoung, et al.
Published: (2026)
UFAL-CUNI at SemEval-2026 Task 11: An Efficient Modular Neuro-symbolic Method for Syllogistic Reasoning
by: Kartáč, Ivan, et al.
Published: (2026)
by: Kartáč, Ivan, et al.
Published: (2026)
Q-STRUM Debate: Query-Driven Contrastive Summarization for Recommendation Comparison
by: Saad, George-Kirollos, et al.
Published: (2025)
by: Saad, George-Kirollos, et al.
Published: (2025)
Text Detoxification as Style Transfer in English and Hindi
by: Mukherjee, Sourabrata, et al.
Published: (2024)
by: Mukherjee, Sourabrata, et al.
Published: (2024)
Teaching LLMs at Charles University: Assignments and Activities
by: Helcl, Jindřich, et al.
Published: (2024)
by: Helcl, Jindřich, et al.
Published: (2024)
How Important is `Perfect' English for Machine Translation Prompts?
by: Schmidtová, Patrícia, et al.
Published: (2025)
by: Schmidtová, Patrícia, et al.
Published: (2025)
Exploring ReAct Prompting for Task-Oriented Dialogue: Insights and Shortcomings
by: Elizabeth, Michelle, et al.
Published: (2024)
by: Elizabeth, Michelle, et al.
Published: (2024)
Multilingual Text Style Transfer: Datasets & Models for Indian Languages
by: Mukherjee, Sourabrata, et al.
Published: (2024)
by: Mukherjee, Sourabrata, et al.
Published: (2024)
Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization?
by: Fu, Xue-Yong, et al.
Published: (2024)
by: Fu, Xue-Yong, et al.
Published: (2024)
Structured Extraction of Real World Medical Knowledge using LLMs for Summarization and Search
by: Kim, Edward, et al.
Published: (2024)
by: Kim, Edward, et al.
Published: (2024)
Intrinsic vs. Extrinsic Evaluation of Czech Sentence Embeddings: Semantic Relevance Doesn't Help with MT Evaluation
by: Barančíková, Petra, et al.
Published: (2025)
by: Barančíková, Petra, et al.
Published: (2025)
We Should Evaluate Real-World Impact
by: Reiter, Ehud
Published: (2025)
by: Reiter, Ehud
Published: (2025)
APPLS: Evaluating Evaluation Metrics for Plain Language Summarization
by: Guo, Yue, et al.
Published: (2023)
by: Guo, Yue, et al.
Published: (2023)
CASPR: Automated Evaluation Metric for Contrastive Summarization
by: Ananthamurugan, Nirupan, et al.
Published: (2024)
by: Ananthamurugan, Nirupan, et al.
Published: (2024)
Calibrating Model-Based Evaluation Metrics for Summarization
by: Liu, Hongye, et al.
Published: (2026)
by: Liu, Hongye, et al.
Published: (2026)
RWESummary: A Framework and Test for Choosing Large Language Models to Summarize Real-World Evidence (RWE) Studies
by: Mukerji, Arjun, et al.
Published: (2025)
by: Mukerji, Arjun, et al.
Published: (2025)
PSentScore: Evaluating Sentiment Polarity in Dialogue Summarization
by: Zhou, Yongxin, et al.
Published: (2023)
by: Zhou, Yongxin, et al.
Published: (2023)
Similar Items
-
Automatic Metrics in Natural Language Generation: A Survey of Current Evaluation Practices
by: Schmidtová, Patrícia, et al.
Published: (2024) -
factgenie: A Framework for Span-based Evaluation of Generated Texts
by: Kasner, Zdeněk, et al.
Published: (2024) -
Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs
by: Balloccu, Simone, et al.
Published: (2024) -
LLMs as Span Annotators: A Comparative Study of LLMs and Humans
by: Kasner, Zdeněk, et al.
Published: (2025) -
Sentence Embeddings as an intermediate target in end-to-end summarisation
by: Zembrzuski, Maciej, et al.
Published: (2025)