What Has Been Lost with Synthetic Evaluation?
Fuente:
arXiv
Saved in:
| Main Authors: | Gill, Alexander, Ravichander, Abhilasha, Marasović, Ana |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
by: Ravichander, Abhilasha, et al.
Published: (2025)
by: Ravichander, Abhilasha, et al.
Published: (2025)
On Evaluating Explanation Utility for Human-AI Decision Making in NLP
by: Chaleshtori, Fateme Hashemi, et al.
Published: (2024)
by: Chaleshtori, Fateme Hashemi, et al.
Published: (2024)
Has It All Been Solved? Open NLP Research Questions Not Solved by Large Language Models
by: Ignat, Oana, et al.
Published: (2023)
by: Ignat, Oana, et al.
Published: (2023)
Chain-of-Thought Unfaithfulness as Disguised Accuracy
by: Bentham, Oliver, et al.
Published: (2024)
by: Bentham, Oliver, et al.
Published: (2024)
Lost-in-the-Middle in Long-Text Generation: Synthetic Dataset, Evaluation Framework, and Mitigation
by: Zhang, Junhao, et al.
Published: (2025)
by: Zhang, Junhao, et al.
Published: (2025)
WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries
by: Zhao, Wenting, et al.
Published: (2024)
by: Zhao, Wenting, et al.
Published: (2024)
Agent Lumos: Unified and Modular Training for Open-Source Language Agents
by: Yin, Da, et al.
Published: (2023)
by: Yin, Da, et al.
Published: (2023)
MacGyver: Are Large Language Models Creative Problem Solvers?
by: Tian, Yufei, et al.
Published: (2023)
by: Tian, Yufei, et al.
Published: (2023)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
by: Lin, Bill Yuchen, et al.
Published: (2024)
by: Lin, Bill Yuchen, et al.
Published: (2024)
Teaching People LLM's Errors and Getting it Right
by: Stringham, Nathan, et al.
Published: (2025)
by: Stringham, Nathan, et al.
Published: (2025)
What Works for 'Lost-in-the-Middle' in LLMs? A Study on GM-Extract and Mitigations
by: Gupte, Mihir, et al.
Published: (2025)
by: Gupte, Mihir, et al.
Published: (2025)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
The Oracle Has Spoken: A Multi-Aspect Evaluation of Dialogue in Pythia
by: Chen, Zixun, et al.
Published: (2025)
by: Chen, Zixun, et al.
Published: (2025)
Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress
by: Proietti, Lorenzo, et al.
Published: (2025)
by: Proietti, Lorenzo, et al.
Published: (2025)
The Term 'Agent' Has Been Diluted Beyond Utility and Requires Redefinition
by: Bent, Brinnae
Published: (2025)
by: Bent, Brinnae
Published: (2025)
Lost in the Source Language: How Large Language Models Evaluate the Quality of Machine Translation
by: Huang, Xu, et al.
Published: (2024)
by: Huang, Xu, et al.
Published: (2024)
Lost in the Logic: An Evaluation of Large Language Models' Reasoning Capabilities on LSAT Logic Games
by: Malik, Saumya
Published: (2024)
by: Malik, Saumya
Published: (2024)
Spanish and LLM Benchmarks: is MMLU Lost in Translation?
by: Plaza, Irene, et al.
Published: (2024)
by: Plaza, Irene, et al.
Published: (2024)
The Art of Saying No: Contextual Noncompliance in Language Models
by: Brahman, Faeze, et al.
Published: (2024)
by: Brahman, Faeze, et al.
Published: (2024)
Blended RAG: Improving RAG (Retriever-Augmented Generation) Accuracy with Semantic Search and Hybrid Query-Based Retrievers
by: Sawarkar, Kunal, et al.
Published: (2024)
by: Sawarkar, Kunal, et al.
Published: (2024)
Can we Evaluate RAGs with Synthetic Data?
by: van Elburg, Jonas, et al.
Published: (2025)
by: van Elburg, Jonas, et al.
Published: (2025)
Creativity Has Left the Chat: The Price of Debiasing Language Models
by: Mohammadi, Behnam
Published: (2024)
by: Mohammadi, Behnam
Published: (2024)
Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
by: Calderon, Nitay, et al.
Published: (2026)
by: Calderon, Nitay, et al.
Published: (2026)
Lost in Stories: Consistency Bugs in Long Story Generation by LLMs
by: Li, Junjie, et al.
Published: (2026)
by: Li, Junjie, et al.
Published: (2026)
Lost in the Noise: How Reasoning Models Fail with Contextual Distractors
by: Lee, Seongyun, et al.
Published: (2026)
by: Lee, Seongyun, et al.
Published: (2026)
Has this Fact been Edited? Detecting Knowledge Edits in Language Models
by: Youssef, Paul, et al.
Published: (2024)
by: Youssef, Paul, et al.
Published: (2024)
Societal AI Research Has Become Less Interdisciplinary
by: Markus, Dror Kris, et al.
Published: (2025)
by: Markus, Dror Kris, et al.
Published: (2025)
Affective Computing Has Changed: The Foundation Model Disruption
by: Schuller, Björn, et al.
Published: (2024)
by: Schuller, Björn, et al.
Published: (2024)
"Lost-in-the-Later": Framework for Quantifying Contextual Grounding in Large Language Models
by: Tao, Yufei, et al.
Published: (2025)
by: Tao, Yufei, et al.
Published: (2025)
Retracing the Past: LLMs Emit Training Data When They Get Lost
by: Ko, Myeongseob, et al.
Published: (2025)
by: Ko, Myeongseob, et al.
Published: (2025)
Lost in Translation? Exploring the Shift in Grammatical Gender from Latin to Occitan
by: Chatterjee, Ahan, et al.
Published: (2026)
by: Chatterjee, Ahan, et al.
Published: (2026)
Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
by: Liu, Geng, et al.
Published: (2026)
by: Liu, Geng, et al.
Published: (2026)
Lost in Translation: Latent Concept Misalignment in Text-to-Image Diffusion Models
by: Zhao, Juntu, et al.
Published: (2024)
by: Zhao, Juntu, et al.
Published: (2024)
A Controlled Synthetic Benchmark for Educational Aspect-Based Sentiment Analysis
by: Aperstein, Yehudit, et al.
Published: (2026)
by: Aperstein, Yehudit, et al.
Published: (2026)
Lost in Translation? A Comparative Study on the Cross-Lingual Transfer of Composite Harms
by: Shukla, Vaibhav, et al.
Published: (2026)
by: Shukla, Vaibhav, et al.
Published: (2026)
Has Your Pretrained Model Improved? A Multi-head Posterior Based Approach
by: Aboagye, Prince, et al.
Published: (2024)
by: Aboagye, Prince, et al.
Published: (2024)
Personalization Increases Affective Alignment but Has Role-Dependent Effects on Epistemic Independence in LLMs
by: Kelley, Sean W., et al.
Published: (2026)
by: Kelley, Sean W., et al.
Published: (2026)
Grounding Synthetic Data Evaluations of Language Models in Unsupervised Document Corpora
by: Majurski, Michael, et al.
Published: (2025)
by: Majurski, Michael, et al.
Published: (2025)
Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models
by: Lameris, Harm, et al.
Published: (2025)
by: Lameris, Harm, et al.
Published: (2025)
Lost in the Pipeline: How Well Do Large Language Models Handle Data Preparation?
by: Spreafico, Matteo, et al.
Published: (2025)
by: Spreafico, Matteo, et al.
Published: (2025)
Similar Items
-
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
by: Ravichander, Abhilasha, et al.
Published: (2025) -
On Evaluating Explanation Utility for Human-AI Decision Making in NLP
by: Chaleshtori, Fateme Hashemi, et al.
Published: (2024) -
Has It All Been Solved? Open NLP Research Questions Not Solved by Large Language Models
by: Ignat, Oana, et al.
Published: (2023) -
Chain-of-Thought Unfaithfulness as Disguised Accuracy
by: Bentham, Oliver, et al.
Published: (2024) -
Lost-in-the-Middle in Long-Text Generation: Synthetic Dataset, Evaluation Framework, and Mitigation
by: Zhang, Junhao, et al.
Published: (2025)