References Matter: Investigating the Impact of Reference Set Variation on Summarization Evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | Casola, Silvia, Liu, Yang Janet, Peng, Siyao, Kraus, Oliver, Gatt, Albert, Plank, Barbara |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Different Tastes of Entities: Investigating Human Label Variation in Named Entity Annotations
por: Peng, Siyao, et al.
Publicado: (2024)
por: Peng, Siyao, et al.
Publicado: (2024)
SteerEval: Inference-time Interventions Strengthen Multilingual Generalization in Neural Summarization Metrics
por: Casola, Silvia, et al.
Publicado: (2026)
por: Casola, Silvia, et al.
Publicado: (2026)
CLIMATELI: Evaluating Entity Linking on Climate Change Data
por: Zhou, Shijia, et al.
Publicado: (2024)
por: Zhou, Shijia, et al.
Publicado: (2024)
EVADE: LLM-Based Explanation Generation and Validation for Error Detection in NLI
por: Zuo, Longfei, et al.
Publicado: (2025)
por: Zuo, Longfei, et al.
Publicado: (2025)
MaiBaam Annotation Guidelines
por: Blaschke, Verena, et al.
Publicado: (2024)
por: Blaschke, Verena, et al.
Publicado: (2024)
VariErr NLI: Separating Annotation Error from Human Label Variation
por: Weber-Genzel, Leon, et al.
Publicado: (2024)
por: Weber-Genzel, Leon, et al.
Publicado: (2024)
Mitigating the Impact of Reference Quality on Evaluation of Summarization Systems with Reference-Free Metrics
por: Gigant, Théo, et al.
Publicado: (2024)
por: Gigant, Théo, et al.
Publicado: (2024)
Threading the Needle: Reweaving Chain-of-Thought Reasoning to Explain Human Label Variation
por: Chen, Beiduo, et al.
Publicado: (2025)
por: Chen, Beiduo, et al.
Publicado: (2025)
LiTEx: A Linguistic Taxonomy of Explanations for Understanding Within-Label Variation in Natural Language Inference
por: Hong, Pingjun, et al.
Publicado: (2025)
por: Hong, Pingjun, et al.
Publicado: (2025)
A Rose by Any Other Name: LLM-Generated Explanations Are Good Proxies for Human Explanations to Collect Label Distributions on NLI
por: Chen, Beiduo, et al.
Publicado: (2024)
por: Chen, Beiduo, et al.
Publicado: (2024)
Synthetic Eggs in Many Baskets: The Impact of Synthetic Data Diversity on LLM Fine-Tuning
por: Schaffelder, Max, et al.
Publicado: (2025)
por: Schaffelder, Max, et al.
Publicado: (2025)
Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set
por: Eichin, Florian, et al.
Publicado: (2025)
por: Eichin, Florian, et al.
Publicado: (2025)
MaiBaam: A Multi-Dialectal Bavarian Universal Dependency Treebank
por: Blaschke, Verena, et al.
Publicado: (2024)
por: Blaschke, Verena, et al.
Publicado: (2024)
MultiClimate: Multimodal Stance Detection on Climate Change Videos
por: Wang, Jiawen, et al.
Publicado: (2024)
por: Wang, Jiawen, et al.
Publicado: (2024)
LeWiDi-2025 at NLPerspectives: Third Edition of the Learning with Disagreements Shared Task
por: Leonardelli, Elisa, et al.
Publicado: (2025)
por: Leonardelli, Elisa, et al.
Publicado: (2025)
Evaluating Large Language Models for Cross-Lingual Retrieval
por: Zuo, Longfei, et al.
Publicado: (2025)
por: Zuo, Longfei, et al.
Publicado: (2025)
BoN Appetit Team at LeWiDi-2025: Best-of-N Test-time Scaling Can Not Stomach Annotation Disagreements (Yet)
por: Ruiz, Tomas, et al.
Publicado: (2025)
por: Ruiz, Tomas, et al.
Publicado: (2025)
Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations
por: Hong, Pingjun, et al.
Publicado: (2025)
por: Hong, Pingjun, et al.
Publicado: (2025)
Summarizing long regulatory documents with a multi-step pipeline
por: Sie, Mika, et al.
Publicado: (2024)
por: Sie, Mika, et al.
Publicado: (2024)
On Learning to Summarize with Large Language Models as References
por: Liu, Yixin, et al.
Publicado: (2023)
por: Liu, Yixin, et al.
Publicado: (2023)
Reason to Rote: Rethinking Memorization in Reasoning
por: Du, Yupei, et al.
Publicado: (2025)
por: Du, Yupei, et al.
Publicado: (2025)
Cross-Dialect Information Retrieval: Information Access in Low-Resource and High-Variance Languages
por: Litschko, Robert, et al.
Publicado: (2024)
por: Litschko, Robert, et al.
Publicado: (2024)
Information Asymmetry across Language Varieties: A Case Study on Cantonese-Mandarin and Bavarian-German QA
por: Pei, Renhao, et al.
Publicado: (2026)
por: Pei, Renhao, et al.
Publicado: (2026)
Redundancy Aware Multi-Reference Based Gainwise Evaluation of Extractive Summarization
por: Akter, Mousumi, et al.
Publicado: (2023)
por: Akter, Mousumi, et al.
Publicado: (2023)
EEVEE: An Easy Annotation Tool for Natural Language Processing
por: Sorensen, Axel, et al.
Publicado: (2024)
por: Sorensen, Axel, et al.
Publicado: (2024)
BanglaSummEval: Reference-Free Factual Consistency Evaluation for Bangla Summarization
por: Rafid, Ahmed, et al.
Publicado: (2026)
por: Rafid, Ahmed, et al.
Publicado: (2026)
Evaluating LLM-Generated Versus Human-Authored Responses in Role-Play Dialogues
por: Lu, Dongxu, et al.
Publicado: (2025)
por: Lu, Dongxu, et al.
Publicado: (2025)
"Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?
por: Chen, Beiduo, et al.
Publicado: (2024)
por: Chen, Beiduo, et al.
Publicado: (2024)
CREAM: Comparison-Based Reference-Free ELO-Ranked Automatic Evaluation for Meeting Summarization
por: Gong, Ziwei, et al.
Publicado: (2024)
por: Gong, Ziwei, et al.
Publicado: (2024)
Information-Theoretic Distillation for Reference-less Summarization
por: Jung, Jaehun, et al.
Publicado: (2024)
por: Jung, Jaehun, et al.
Publicado: (2024)
Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential Communication
por: Merlo, Filippo, et al.
Publicado: (2025)
por: Merlo, Filippo, et al.
Publicado: (2025)
Morphological Analysis for the Maltese Language: The Challenges of a Hybrid System
por: Borg, Claudia, et al.
Publicado: (2017)
por: Borg, Claudia, et al.
Publicado: (2017)
A Systematic Analysis of Large Language Models as Soft Reasoners: The Case of Syllogistic Inferences
por: Bertolazzi, Leonardo, et al.
Publicado: (2024)
por: Bertolazzi, Leonardo, et al.
Publicado: (2024)
CV-Probes: Studying the interplay of lexical and world knowledge in visually grounded verb understanding
por: Beňová, Ivana, et al.
Publicado: (2024)
por: Beňová, Ivana, et al.
Publicado: (2024)
Probing Omissions and Distortions in Transformer-based RDF-to-Text Models
por: Faille, Juliette, et al.
Publicado: (2024)
por: Faille, Juliette, et al.
Publicado: (2024)
Sebastian, Basti, Wastl?! Recognizing Named Entities in Bavarian Dialectal Data
por: Peng, Siyao, et al.
Publicado: (2024)
por: Peng, Siyao, et al.
Publicado: (2024)
ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization
por: Bae, Suyoung, et al.
Publicado: (2026)
por: Bae, Suyoung, et al.
Publicado: (2026)
References Indeed Matter? Reference-Free Preference Optimization for Conversational Query Reformulation
por: Kim, Doyoung, et al.
Publicado: (2025)
por: Kim, Doyoung, et al.
Publicado: (2025)
Interpreting Predictive Probabilities: Model Confidence or Human Label Variation?
por: Baan, Joris, et al.
Publicado: (2024)
por: Baan, Joris, et al.
Publicado: (2024)
VAQUUM: Are Vague Quantifiers Grounded in Visual Data?
por: Wong, Hugh Mee, et al.
Publicado: (2025)
por: Wong, Hugh Mee, et al.
Publicado: (2025)
Ejemplares similares
-
Different Tastes of Entities: Investigating Human Label Variation in Named Entity Annotations
por: Peng, Siyao, et al.
Publicado: (2024) -
SteerEval: Inference-time Interventions Strengthen Multilingual Generalization in Neural Summarization Metrics
por: Casola, Silvia, et al.
Publicado: (2026) -
CLIMATELI: Evaluating Entity Linking on Climate Change Data
por: Zhou, Shijia, et al.
Publicado: (2024) -
EVADE: LLM-Based Explanation Generation and Validation for Error Detection in NLI
por: Zuo, Longfei, et al.
Publicado: (2025) -
MaiBaam Annotation Guidelines
por: Blaschke, Verena, et al.
Publicado: (2024)