Evaluating the Evaluators: Are readability metrics good measures of readability?
Fuente:
arXiv
Saved in:
| Main Authors: | Cachola, Isabel, Khashabi, Daniel, Dredze, Mark |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Making FETCH! Happen: Finding Emergent Dog Whistles Through Common Habitats
by: Sasse, Kuleen, et al.
Published: (2024)
by: Sasse, Kuleen, et al.
Published: (2024)
Exploring the change in scientific readability following the release of ChatGPT
by: Alsudais, Abdulkareem
Published: (2025)
by: Alsudais, Abdulkareem
Published: (2025)
A study of Vietnamese readability assessing through semantic and statistical features
by: Le, Hung Tuan, et al.
Published: (2024)
by: Le, Hung Tuan, et al.
Published: (2024)
ReadCtrl: Personalizing text generation with readability-controlled instruction learning
by: Tran, Hieu, et al.
Published: (2024)
by: Tran, Hieu, et al.
Published: (2024)
MedReadCtrl: Personalizing medical text generation with readability-controlled instruction learning
by: Tran, Hieu, et al.
Published: (2025)
by: Tran, Hieu, et al.
Published: (2025)
Challenging the Evaluator: LLM Sycophancy Under User Rebuttal
by: Kim, Sungwon, et al.
Published: (2025)
by: Kim, Sungwon, et al.
Published: (2025)
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
by: Jahara, Fatima, et al.
Published: (2025)
by: Jahara, Fatima, et al.
Published: (2025)
What makes a good metric? Evaluating automatic metrics for text-to-image consistency
by: Ross, Candace, et al.
Published: (2024)
by: Ross, Candace, et al.
Published: (2024)
MedScore: Generalizable Factuality Evaluation of Free-Form Medical Answers by Domain-adapted Claim Decomposition and Verification
by: Huang, Heyuan, et al.
Published: (2025)
by: Huang, Heyuan, et al.
Published: (2025)
Evaluating Biases in Context-Dependent Health Questions
by: Levy, Sharon, et al.
Published: (2024)
by: Levy, Sharon, et al.
Published: (2024)
Amuro and Char: Analyzing the Relationship between Pre-Training and Fine-Tuning of Large Language Models
by: Sun, Kaiser, et al.
Published: (2024)
by: Sun, Kaiser, et al.
Published: (2024)
Hell or High Water: Evaluating Agentic Recovery from External Failures
by: Wang, Andrew, et al.
Published: (2025)
by: Wang, Andrew, et al.
Published: (2025)
Can one size fit all?: Measuring Failure in Multi-Document Summarization Domain Transfer
by: DeLucia, Alexandra, et al.
Published: (2025)
by: DeLucia, Alexandra, et al.
Published: (2025)
RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models
by: An, Bang, et al.
Published: (2025)
by: An, Bang, et al.
Published: (2025)
RORA: Robust Free-Text Rationale Evaluation
by: Jiang, Zhengping, et al.
Published: (2024)
by: Jiang, Zhengping, et al.
Published: (2024)
GOLD PANNING: Strategic Context Shuffling for Needle-in-Haystack Reasoning
by: Byerly, Adam, et al.
Published: (2025)
by: Byerly, Adam, et al.
Published: (2025)
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning
by: Li, Tianjian, et al.
Published: (2025)
by: Li, Tianjian, et al.
Published: (2025)
Self-Consistency Falls Short! The Adverse Effects of Positional Bias on Long-Context Problems
by: Byerly, Adam, et al.
Published: (2024)
by: Byerly, Adam, et al.
Published: (2024)
DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generation
by: Wanner, Miriam, et al.
Published: (2024)
by: Wanner, Miriam, et al.
Published: (2024)
arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation
by: Wang, Weiqi, et al.
Published: (2025)
by: Wang, Weiqi, et al.
Published: (2025)
Task Matters: Knowledge Requirements Shape LLM Responses to Context-Memory Conflict
by: Sun, Kaiser, et al.
Published: (2025)
by: Sun, Kaiser, et al.
Published: (2025)
Knowledge-Centric Templatic Views of Documents
by: Cachola, Isabel, et al.
Published: (2024)
by: Cachola, Isabel, et al.
Published: (2024)
LLMs are Better Than You Think: Label-Guided In-Context Learning for Named Entity Recognition
by: Bai, Fan, et al.
Published: (2025)
by: Bai, Fan, et al.
Published: (2025)
Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions
by: Chen, Hanjie, et al.
Published: (2024)
by: Chen, Hanjie, et al.
Published: (2024)
Are Clinical T5 Models Better for Clinical Text?
by: Li, Yahan, et al.
Published: (2024)
by: Li, Yahan, et al.
Published: (2024)
Local Compositional Complexity: How to Detect a Human-readable Messsage
by: Mahon, Louis
Published: (2025)
by: Mahon, Louis
Published: (2025)
A Standardized Machine-readable Dataset Documentation Format for Responsible AI
by: Jain, Nitisha, et al.
Published: (2024)
by: Jain, Nitisha, et al.
Published: (2024)
Open Datasheets: Machine-readable Documentation for Open Datasets and Responsible AI Assessments
by: Roman, Anthony Cintron, et al.
Published: (2023)
by: Roman, Anthony Cintron, et al.
Published: (2023)
The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
by: Uzunoglu, Arda, et al.
Published: (2025)
by: Uzunoglu, Arda, et al.
Published: (2025)
Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
by: Uzunoglu, Arda, et al.
Published: (2026)
by: Uzunoglu, Arda, et al.
Published: (2026)
On the Failure of Latent State Persistence in Large Language Models
by: Huang, Jen-tse, et al.
Published: (2025)
by: Huang, Jen-tse, et al.
Published: (2025)
Punctuation as readability and textuality factor in technical discourse
by: Carmen Sancho Guinda
Published: (2002)
by: Carmen Sancho Guinda
Published: (2002)
A Closer Look at Claim Decomposition
by: Wanner, Miriam, et al.
Published: (2024)
by: Wanner, Miriam, et al.
Published: (2024)
*-PLUIE: Personalisable metric with Llm Used for Improved Evaluation
by: Lemesle, Quentin, et al.
Published: (2026)
by: Lemesle, Quentin, et al.
Published: (2026)
Give me Some Hard Questions: Synthetic Data Generation for Clinical QA
by: Bai, Fan, et al.
Published: (2024)
by: Bai, Fan, et al.
Published: (2024)
Schema-Driven Information Extraction from Heterogeneous Tables
by: Bai, Fan, et al.
Published: (2023)
by: Bai, Fan, et al.
Published: (2023)
IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
by: Mishra, Aayush, et al.
Published: (2025)
by: Mishra, Aayush, et al.
Published: (2025)
Do pretrained Transformers Learn In-Context by Gradient Descent?
by: Shen, Lingfeng, et al.
Published: (2023)
by: Shen, Lingfeng, et al.
Published: (2023)
WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment
by: Ou, Jiefu, et al.
Published: (2024)
by: Ou, Jiefu, et al.
Published: (2024)
Are Finer Citations Always Better? Rethinking Granularity for Attributed Generation
by: Wang, Hexuan, et al.
Published: (2026)
by: Wang, Hexuan, et al.
Published: (2026)
Similar Items
-
Making FETCH! Happen: Finding Emergent Dog Whistles Through Common Habitats
by: Sasse, Kuleen, et al.
Published: (2024) -
Exploring the change in scientific readability following the release of ChatGPT
by: Alsudais, Abdulkareem
Published: (2025) -
A study of Vietnamese readability assessing through semantic and statistical features
by: Le, Hung Tuan, et al.
Published: (2024) -
ReadCtrl: Personalizing text generation with readability-controlled instruction learning
by: Tran, Hieu, et al.
Published: (2024) -
MedReadCtrl: Personalizing medical text generation with readability-controlled instruction learning
by: Tran, Hieu, et al.
Published: (2025)