Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?
Fuente:
arXiv
Saved in:
| Main Authors: | Barnes, Jeremy, Perez, Naiara, Bonet-Jover, Alba, Altuna, Begoña |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NoticIA: A Clickbait Article Summarization Dataset in Spanish
by: García-Ferrero, Iker, et al.
Published: (2024)
by: García-Ferrero, Iker, et al.
Published: (2024)
EuskañolDS: A Naturally Sourced Corpus for Basque-Spanish Code-Switching
by: Heredia, Maite, et al.
Published: (2025)
by: Heredia, Maite, et al.
Published: (2025)
Automatic Essay Scoring and Feedback Generation in Basque Language Learning
by: Azurmendi, Ekhi, et al.
Published: (2025)
by: Azurmendi, Ekhi, et al.
Published: (2025)
XNLIeu: a dataset for cross-lingual NLI in Basque
by: Heredia, Maite, et al.
Published: (2024)
by: Heredia, Maite, et al.
Published: (2024)
What's under the hood: Investigating Automatic Metrics on Meeting Summarization
by: Kirstein, Frederic, et al.
Published: (2024)
by: Kirstein, Frederic, et al.
Published: (2024)
When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation
by: Sun, Bian, et al.
Published: (2026)
by: Sun, Bian, et al.
Published: (2026)
Evaluating Metrics for Safety with LLM-as-Judges
by: Clegg, Kester, et al.
Published: (2025)
by: Clegg, Kester, et al.
Published: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
by: Yang, Langqi, et al.
Published: (2025)
by: Yang, Langqi, et al.
Published: (2025)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
Latxa: An Open Language Model and Evaluation Suite for Basque
by: Etxaniz, Julen, et al.
Published: (2024)
by: Etxaniz, Julen, et al.
Published: (2024)
Automatic Summarization of Long Documents
by: Chhibbar, Naman, et al.
Published: (2024)
by: Chhibbar, Naman, et al.
Published: (2024)
Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge
by: Fujinuma, Yoshinari
Published: (2025)
by: Fujinuma, Yoshinari
Published: (2025)
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
by: Ramprasad, Sanjana, et al.
Published: (2024)
by: Ramprasad, Sanjana, et al.
Published: (2024)
SemBench: A Universal Semantic Framework for LLM Evaluation
by: Zubillaga, Mikel, et al.
Published: (2026)
by: Zubillaga, Mikel, et al.
Published: (2026)
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
by: Han, Steve, et al.
Published: (2025)
by: Han, Steve, et al.
Published: (2025)
Improve LLM-based Automatic Essay Scoring with Linguistic Features
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
Knowledge Distillation of LLM for Automatic Scoring of Science Education Assessments
by: Latif, Ehsan, et al.
Published: (2023)
by: Latif, Ehsan, et al.
Published: (2023)
AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
by: Ryan, Michael J., et al.
Published: (2025)
by: Ryan, Michael J., et al.
Published: (2025)
CNsum:Automatic Summarization for Chinese News Text
by: Zhao, Yu, et al.
Published: (2025)
by: Zhao, Yu, et al.
Published: (2025)
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
by: Kim, Jane Paik
Published: (2026)
by: Kim, Jane Paik
Published: (2026)
DecMetrics: Structured Claim Decomposition Scoring for Factually Consistent LLM Outputs
by: Huang, Minghui
Published: (2025)
by: Huang, Minghui
Published: (2025)
Comparing Approaches to Automatic Summarization in Less-Resourced Languages
by: Palen-Michel, Chester, et al.
Published: (2025)
by: Palen-Michel, Chester, et al.
Published: (2025)
Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation
by: Lin, Wei-Hsiang, et al.
Published: (2025)
by: Lin, Wei-Hsiang, et al.
Published: (2025)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
by: Shi, Lin, et al.
Published: (2024)
by: Shi, Lin, et al.
Published: (2024)
Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy
by: Thompson, Brian, et al.
Published: (2024)
by: Thompson, Brian, et al.
Published: (2024)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
by: Hong, Yihan, et al.
Published: (2026)
by: Hong, Yihan, et al.
Published: (2026)
When LLM Judge Scores Look Good but Best-of-N Decisions Fail
by: Landesberg, Eddie
Published: (2026)
by: Landesberg, Eddie
Published: (2026)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
by: Alam, Firoj, et al.
Published: (2026)
by: Alam, Firoj, et al.
Published: (2026)
Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
by: He, Zhonghao, et al.
Published: (2025)
by: He, Zhonghao, et al.
Published: (2025)
LeMAJ (Legal LLM-as-a-Judge): Bridging Legal Reasoning and LLM Evaluation
by: Enguehard, Joseph, et al.
Published: (2025)
by: Enguehard, Joseph, et al.
Published: (2025)
Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge
by: Sun, Xin, et al.
Published: (2026)
by: Sun, Xin, et al.
Published: (2026)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
by: Wang, Yidong, et al.
Published: (2025)
by: Wang, Yidong, et al.
Published: (2025)
Learning to Summarize from LLM-generated Feedback
by: Song, Hwanjun, et al.
Published: (2024)
by: Song, Hwanjun, et al.
Published: (2024)
A Survey on LLM-as-a-Judge
by: Gu, Jiawei, et al.
Published: (2024)
by: Gu, Jiawei, et al.
Published: (2024)
Do LLMs Judge Distantly Supervised Named Entity Labels Well? Constructing the JudgeWEL Dataset
by: Plum, Alistair, et al.
Published: (2026)
by: Plum, Alistair, et al.
Published: (2026)
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
by: Collot, Stephane, et al.
Published: (2025)
by: Collot, Stephane, et al.
Published: (2025)
Similar Data Points Identification with LLM: A Human-in-the-loop Strategy Using Summarization and Hidden State Insights
by: Zeng, Xianlong, et al.
Published: (2024)
by: Zeng, Xianlong, et al.
Published: (2024)
Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics
by: Cho, Yousang, et al.
Published: (2025)
by: Cho, Yousang, et al.
Published: (2025)
Beyond Overlap Metrics: Rewarding Reasoning and Preferences for Faithful Multi-Role Dialogue Summarization
by: Mei, Xiaoyong, et al.
Published: (2026)
by: Mei, Xiaoyong, et al.
Published: (2026)
A LLM-Powered Automatic Grading Framework with Human-Level Guidelines Optimization
by: Chu, Yucheng, et al.
Published: (2024)
by: Chu, Yucheng, et al.
Published: (2024)
Similar Items
-
NoticIA: A Clickbait Article Summarization Dataset in Spanish
by: García-Ferrero, Iker, et al.
Published: (2024) -
EuskañolDS: A Naturally Sourced Corpus for Basque-Spanish Code-Switching
by: Heredia, Maite, et al.
Published: (2025) -
Automatic Essay Scoring and Feedback Generation in Basque Language Learning
by: Azurmendi, Ekhi, et al.
Published: (2025) -
XNLIeu: a dataset for cross-lingual NLI in Basque
by: Heredia, Maite, et al.
Published: (2024) -
What's under the hood: Investigating Automatic Metrics on Meeting Summarization
by: Kirstein, Frederic, et al.
Published: (2024)