Don't Sweat the Small Stuff: Segment-Level Meta-Evaluation Based on Pairwise Difference Correlation
Fuente:
arXiv
Salvato in:
| Autori principali: | DiIanni, Colten, Deutsch, Daniel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
di: Riley, Parker, et al.
Pubblicazione: (2025)
di: Riley, Parker, et al.
Pubblicazione: (2025)
Why Don't Prompt-Based Fairness Metrics Correlate?
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2024)
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2024)
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
di: Gur-Arieh, Yoav, et al.
Pubblicazione: (2026)
di: Gur-Arieh, Yoav, et al.
Pubblicazione: (2026)
Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy
di: Thompson, Brian, et al.
Pubblicazione: (2024)
di: Thompson, Brian, et al.
Pubblicazione: (2024)
Reasoning Models Don't Just Think Longer, They Move Differently
di: Gjølbye, Anders, et al.
Pubblicazione: (2026)
di: Gjølbye, Anders, et al.
Pubblicazione: (2026)
Hatevolution: What Static Benchmarks Don't Tell Us
di: Di Bonaventura, Chiara, et al.
Pubblicazione: (2025)
di: Di Bonaventura, Chiara, et al.
Pubblicazione: (2025)
Don't Adapt Small Language Models for Tools; Adapt Tool Schemas to the Models
di: Lee, Jonggeun, et al.
Pubblicazione: (2025)
di: Lee, Jonggeun, et al.
Pubblicazione: (2025)
Don't Touch My Diacritics
di: Gorman, Kyle, et al.
Pubblicazione: (2024)
di: Gorman, Kyle, et al.
Pubblicazione: (2024)
Don't Pay Attention
di: Hammoud, Mohammad, et al.
Pubblicazione: (2025)
di: Hammoud, Mohammad, et al.
Pubblicazione: (2025)
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
di: Hernandez, Adriano
Pubblicazione: (2024)
di: Hernandez, Adriano
Pubblicazione: (2024)
Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks
di: Lumer, Elias, et al.
Pubblicazione: (2026)
di: Lumer, Elias, et al.
Pubblicazione: (2026)
Don't Throw Away Your Pretrained Model
di: Feng, Shangbin, et al.
Pubblicazione: (2025)
di: Feng, Shangbin, et al.
Pubblicazione: (2025)
Don't Say No: Jailbreaking LLM by Suppressing Refusal
di: Zhou, Yukai, et al.
Pubblicazione: (2024)
di: Zhou, Yukai, et al.
Pubblicazione: (2024)
Don't Take the Premise for Granted: Evaluating the Premise Critique Ability of Large Language Models
di: Li, Jinzhe, et al.
Pubblicazione: (2025)
di: Li, Jinzhe, et al.
Pubblicazione: (2025)
Why Don't You Know? Evaluating the Impact of Uncertainty Sources on Uncertainty Quantification in LLMs
di: Goloburda, Maiya, et al.
Pubblicazione: (2026)
di: Goloburda, Maiya, et al.
Pubblicazione: (2026)
Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell
di: Lu, Taiming, et al.
Pubblicazione: (2024)
di: Lu, Taiming, et al.
Pubblicazione: (2024)
Think, But Don't Overthink: Reproducing Recursive Language Models
di: Wang, Daren
Pubblicazione: (2026)
di: Wang, Daren
Pubblicazione: (2026)
Predict, Don't React: Value-Based Safety Forecasting for LLM Streaming
di: Kavumba, Pride, et al.
Pubblicazione: (2026)
di: Kavumba, Pride, et al.
Pubblicazione: (2026)
Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty
di: Ren, Jingyi, et al.
Pubblicazione: (2026)
di: Ren, Jingyi, et al.
Pubblicazione: (2026)
Reasoning Models Reason Well, Until They Don't
di: Rameshkumar, Revanth, et al.
Pubblicazione: (2025)
di: Rameshkumar, Revanth, et al.
Pubblicazione: (2025)
Don't Throw Away Data: Better Sequence Knowledge Distillation
di: Wang, Jun, et al.
Pubblicazione: (2024)
di: Wang, Jun, et al.
Pubblicazione: (2024)
Don't Command, Cultivate: An Exploratory Study of System-2 Alignment
di: Wang, Yuhang, et al.
Pubblicazione: (2024)
di: Wang, Yuhang, et al.
Pubblicazione: (2024)
LLM Cyber Evaluations Don't Capture Real-World Risk
di: Lukošiūtė, Kamilė, et al.
Pubblicazione: (2025)
di: Lukošiūtė, Kamilė, et al.
Pubblicazione: (2025)
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
di: Moon, Jiwon, et al.
Pubblicazione: (2025)
di: Moon, Jiwon, et al.
Pubblicazione: (2025)
Honest AI: Fine-Tuning "Small" Language Models to Say "I Don't Know", and Reducing Hallucination in RAG
di: Chen, Xinxi, et al.
Pubblicazione: (2024)
di: Chen, Xinxi, et al.
Pubblicazione: (2024)
Don't Walk the Line: Boundary Guidance for Filtered Generation
di: Ball, Sarah, et al.
Pubblicazione: (2025)
di: Ball, Sarah, et al.
Pubblicazione: (2025)
Language Models Don't Learn the Physical Manifestation of Language
di: Lee, Bruce W., et al.
Pubblicazione: (2024)
di: Lee, Bruce W., et al.
Pubblicazione: (2024)
Can AI Assistants Know What They Don't Know?
di: Cheng, Qinyuan, et al.
Pubblicazione: (2024)
di: Cheng, Qinyuan, et al.
Pubblicazione: (2024)
Frictional Agent Alignment Framework: Slow Down and Don't Break Things
di: Nath, Abhijnan, et al.
Pubblicazione: (2025)
di: Nath, Abhijnan, et al.
Pubblicazione: (2025)
Don't Lose Focus: Activation Steering via Key-Orthogonal Projections
di: Luo, Haoyan, et al.
Pubblicazione: (2026)
di: Luo, Haoyan, et al.
Pubblicazione: (2026)
Pointer-Generator Networks for Low-Resource Machine Translation: Don't Copy That!
di: Bafna, Niyati, et al.
Pubblicazione: (2024)
di: Bafna, Niyati, et al.
Pubblicazione: (2024)
Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models
di: Parmar, Jupinder, et al.
Pubblicazione: (2024)
di: Parmar, Jupinder, et al.
Pubblicazione: (2024)
Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs
di: Hans, Abhimanyu, et al.
Pubblicazione: (2024)
di: Hans, Abhimanyu, et al.
Pubblicazione: (2024)
Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
di: Balepur, Nishant, et al.
Pubblicazione: (2026)
di: Balepur, Nishant, et al.
Pubblicazione: (2026)
Show, Don't Tell: Evaluating Large Language Models Beyond Textual Understanding with ChildPlay
di: de Carvalho, Gonçalo Hora, et al.
Pubblicazione: (2024)
di: de Carvalho, Gonçalo Hora, et al.
Pubblicazione: (2024)
Bridging and Modeling Correlations in Pairwise Data for Direct Preference Optimization
di: Jiang, Yuxin, et al.
Pubblicazione: (2024)
di: Jiang, Yuxin, et al.
Pubblicazione: (2024)
Enhancing Human Evaluation in Machine Translation with Comparative Judgment
di: Song, Yixiao, et al.
Pubblicazione: (2025)
di: Song, Yixiao, et al.
Pubblicazione: (2025)
Analyzing and Evaluating Correlation Measures in NLG Meta-Evaluation
di: Gao, Mingqi, et al.
Pubblicazione: (2024)
di: Gao, Mingqi, et al.
Pubblicazione: (2024)
Don't Erase, Inform! Detecting and Contextualizing Harmful Language in Cultural Heritage Collections
di: Mastromichalakis, Orfeas Menis, et al.
Pubblicazione: (2025)
di: Mastromichalakis, Orfeas Menis, et al.
Pubblicazione: (2025)
Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding
di: Ignatev, Daniil, et al.
Pubblicazione: (2025)
di: Ignatev, Daniil, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
di: Riley, Parker, et al.
Pubblicazione: (2025) -
Why Don't Prompt-Based Fairness Metrics Correlate?
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2024) -
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
di: Gur-Arieh, Yoav, et al.
Pubblicazione: (2026) -
Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy
di: Thompson, Brian, et al.
Pubblicazione: (2024) -
Reasoning Models Don't Just Think Longer, They Move Differently
di: Gjølbye, Anders, et al.
Pubblicazione: (2026)