Penalizing Length: Uncovering Systematic Bias in Quality Estimation Metrics
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yilin, Xu, Wenda, Liu, Zhongtao, Nakagawa, Tetsuji, Freitag, Markus |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task
by: Juraska, Juraj, et al.
Published: (2025)
by: Juraska, Juraj, et al.
Published: (2025)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
by: Xu, Wenda, et al.
Published: (2025)
by: Xu, Wenda, et al.
Published: (2025)
Mitigating Metric Bias in Minimum Bayes Risk Decoding
by: Kovacs, Geza, et al.
Published: (2024)
by: Kovacs, Geza, et al.
Published: (2024)
On the Implications of Verbose LLM Outputs: A Case Study in Translation Evaluation
by: Briakou, Eleftheria, et al.
Published: (2024)
by: Briakou, Eleftheria, et al.
Published: (2024)
LLMRefine: Pinpointing and Refining Large Language Models via Fine-Grained Actionable Feedback
by: Xu, Wenda, et al.
Published: (2023)
by: Xu, Wenda, et al.
Published: (2023)
Beyond Human-Only: Evaluating Human-Machine Collaboration for Collecting High-Quality Translation Data
by: Liu, Zhongtao, et al.
Published: (2024)
by: Liu, Zhongtao, et al.
Published: (2024)
MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task
by: Juraska, Juraj, et al.
Published: (2024)
by: Juraska, Juraj, et al.
Published: (2024)
An Automated Length-Aware Quality Metric for Summarization
by: Foland, Andrew D.
Published: (2025)
by: Foland, Andrew D.
Published: (2025)
Generating Difficult-to-Translate Texts
by: Zouhar, Vilém, et al.
Published: (2025)
by: Zouhar, Vilém, et al.
Published: (2025)
GAMBIT+: A Challenge Set for Evaluating Gender Bias in Machine Translation Quality Estimation Metrics
by: Filandrianos, Giorgos, et al.
Published: (2025)
by: Filandrianos, Giorgos, et al.
Published: (2025)
Introducing the NewsPaLM MBR and QE Dataset: LLM-Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data
by: Finkelstein, Mara, et al.
Published: (2024)
by: Finkelstein, Mara, et al.
Published: (2024)
Quality-Aware Translation Models: Efficient Generation and Quality Estimation in a Single Model
by: Tomani, Christian, et al.
Published: (2023)
by: Tomani, Christian, et al.
Published: (2023)
Translating Step-by-Step: Decomposing the Translation Process for Improved Translation Quality of Long-Form Texts
by: Briakou, Eleftheria, et al.
Published: (2024)
by: Briakou, Eleftheria, et al.
Published: (2024)
Uncovering Factor Level Preferences to Improve Human-Model Alignment
by: Oh, Juhyun, et al.
Published: (2024)
by: Oh, Juhyun, et al.
Published: (2024)
Meta-aware Learning in text-to-SQL Large Language Model
by: Zhang, Wenda
Published: (2025)
by: Zhang, Wenda
Published: (2025)
Stepwise Penalization for Length-Efficient Chain-of-Thought Reasoning
by: Li, Xintong, et al.
Published: (2026)
by: Li, Xintong, et al.
Published: (2026)
SPAGBias: Uncovering and Tracing Structured Spatial Gender Bias in Large Language Models
by: Su, Binxian, et al.
Published: (2026)
by: Su, Binxian, et al.
Published: (2026)
Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement
by: Xu, Wenda, et al.
Published: (2024)
by: Xu, Wenda, et al.
Published: (2024)
MALIBU Benchmark: Multi-Agent LLM Implicit Bias Uncovered
by: Mirza, Imran, et al.
Published: (2025)
by: Mirza, Imran, et al.
Published: (2025)
Evaluating Metrics for Bias in Word Embeddings
by: Schröder, Sarah, et al.
Published: (2021)
by: Schröder, Sarah, et al.
Published: (2021)
Mitigating Label Length Bias in Large Language Models
by: Sanz-Guerrero, Mario, et al.
Published: (2025)
by: Sanz-Guerrero, Mario, et al.
Published: (2025)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
by: Demchak, Nathaniel, et al.
Published: (2024)
by: Demchak, Nathaniel, et al.
Published: (2024)
Explaining Length Bias in LLM-Based Preference Evaluations
by: Hu, Zhengyu, et al.
Published: (2024)
by: Hu, Zhengyu, et al.
Published: (2024)
Towards Region-aware Bias Evaluation Metrics
by: Borah, Angana, et al.
Published: (2024)
by: Borah, Angana, et al.
Published: (2024)
Veracity Bias and Beyond: Uncovering LLMs' Hidden Beliefs in Problem-Solving Reasoning
by: Zhou, Yue, et al.
Published: (2025)
by: Zhou, Yue, et al.
Published: (2025)
Textual Similarity as a Key Metric in Machine Translation Quality Estimation
by: Sun, Kun, et al.
Published: (2024)
by: Sun, Kun, et al.
Published: (2024)
Data Quality Enhancement on the Basis of Diversity with Large Language Models for Text Classification: Uncovered, Difficult, and Noisy
by: Zeng, Min, et al.
Published: (2024)
by: Zeng, Min, et al.
Published: (2024)
An Empirical Study on the Characteristics of Bias upon Context Length Variation for Bangla
by: Sadhu, Jayanta, et al.
Published: (2024)
by: Sadhu, Jayanta, et al.
Published: (2024)
When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method
by: Zhang, Biao, et al.
Published: (2024)
by: Zhang, Biao, et al.
Published: (2024)
Read Between the Lines: A Benchmark for Uncovering Political Bias in Bangla News Articles
by: Lia, Nusrat Jahan, et al.
Published: (2025)
by: Lia, Nusrat Jahan, et al.
Published: (2025)
Improving Multimodal Contrastive Learning of Sentence Embeddings with Object-Phrase Alignment
by: Zhao, Kaiyan, et al.
Published: (2025)
by: Zhao, Kaiyan, et al.
Published: (2025)
GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression
by: Miao, Zhongtao, et al.
Published: (2026)
by: Miao, Zhongtao, et al.
Published: (2026)
Enhancing Human Evaluation in Machine Translation with Comparative Judgment
by: Song, Yixiao, et al.
Published: (2025)
by: Song, Yixiao, et al.
Published: (2025)
Uncovering Implicit Bias in Large Language Models with Concept Learning Dataset
by: Wang, Leroy Z.
Published: (2025)
by: Wang, Leroy Z.
Published: (2025)
Bias Attribution in Filipino Language Models: Extending a Bias Interpretability Metric for Application on Agglutinative Languages
by: Gamboa, Lance Calvin Lim, et al.
Published: (2025)
by: Gamboa, Lance Calvin Lim, et al.
Published: (2025)
Mitigating Length Bias in RLHF through a Causal Lens
by: Kim, Hyeonji, et al.
Published: (2025)
by: Kim, Hyeonji, et al.
Published: (2025)
More Thinking, More Bias: Length-Driven Position Bias in Reasoning Models
by: Wang, Xiao
Published: (2026)
by: Wang, Xiao
Published: (2026)
You Cannot Feed Two Birds with One Score: the Accuracy-Naturalness Tradeoff in Translation
by: Flamich, Gergely, et al.
Published: (2025)
by: Flamich, Gergely, et al.
Published: (2025)
Machine-Generated Text Localization
by: Zhang, Zhongping, et al.
Published: (2024)
by: Zhang, Zhongping, et al.
Published: (2024)
An Automatic Quality Metric for Evaluating Simultaneous Interpretation
by: Makinae, Mana, et al.
Published: (2024)
by: Makinae, Mana, et al.
Published: (2024)
Similar Items
-
MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task
by: Juraska, Juraj, et al.
Published: (2025) -
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
by: Xu, Wenda, et al.
Published: (2025) -
Mitigating Metric Bias in Minimum Bayes Risk Decoding
by: Kovacs, Geza, et al.
Published: (2024) -
On the Implications of Verbose LLM Outputs: A Case Study in Translation Evaluation
by: Briakou, Eleftheria, et al.
Published: (2024) -
LLMRefine: Pinpointing and Refining Large Language Models via Fine-Grained Actionable Feedback
by: Xu, Wenda, et al.
Published: (2023)