The illusion of a perfect metric: Why evaluating AI's words is harder than it looks
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Oliva, Maria Paz, Correia, Adriana, Vankov, Ivan, Botev, Viktor |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ConSens: Assessing context grounding in open-book question answering
par: Vankov, Ivan, et autres
Publié: (2025)
par: Vankov, Ivan, et autres
Publié: (2025)
The Boy Who Survived: Removing Harry Potter from an LLM is harder than reported
par: Shostack, Adam
Publié: (2024)
par: Shostack, Adam
Publié: (2024)
Evaluating Embedding Frameworks for Scientific Domain
par: Ahmed, Nouman, et autres
Publié: (2025)
par: Ahmed, Nouman, et autres
Publié: (2025)
A word association network methodology for evaluating implicit biases in LLMs compared to humans
par: Abramski, Katherine, et autres
Publié: (2025)
par: Abramski, Katherine, et autres
Publié: (2025)
Automated alignment is harder than you think
par: Bowkis, Aleksandr, et autres
Publié: (2026)
par: Bowkis, Aleksandr, et autres
Publié: (2026)
Faithfulness metric fusion: Improving the evaluation of LLM trustworthiness across domains
par: Malin, Ben, et autres
Publié: (2025)
par: Malin, Ben, et autres
Publié: (2025)
"They parted illusions -- they parted disclaim marinade": Misalignment as structural fidelity in LLMs
par: Costa, Mariana Lins
Publié: (2025)
par: Costa, Mariana Lins
Publié: (2025)
LLMs as annotators of credibility assessment in Danish asylum decisions: evaluating classification performance and errors beyond aggregated metrics
par: Humblot-Renaux, Galadrielle, et autres
Publié: (2026)
par: Humblot-Renaux, Galadrielle, et autres
Publié: (2026)
How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives
par: Ichmoukhamedov, Timour, et autres
Publié: (2024)
par: Ichmoukhamedov, Timour, et autres
Publié: (2024)
LiTransProQA: an LLM-based Literary Translation evaluation metric with Professional Question Answering
par: Zhang, Ran, et autres
Publié: (2025)
par: Zhang, Ran, et autres
Publié: (2025)
System 2 thinking in OpenAI's o1-preview model: Near-perfect performance on a mathematics exam
par: de Winter, Joost, et autres
Publié: (2024)
par: de Winter, Joost, et autres
Publié: (2024)
Why AI-Generated Text Detection Fails: Evidence from Explainable AI Beyond Benchmark Accuracy
par: Pudasaini, Shushanta, et autres
Publié: (2026)
par: Pudasaini, Shushanta, et autres
Publié: (2026)
Investigating the structure of emotions by analyzing similarity and association of emotion words
par: Iwaki, Fumitaka, et autres
Publié: (2026)
par: Iwaki, Fumitaka, et autres
Publié: (2026)
LongTail-Swap: benchmarking language models' abilities on rare words
par: Algayres, Robin, et autres
Publié: (2025)
par: Algayres, Robin, et autres
Publié: (2025)
Playing with words: Comparing the vocabulary and lexical diversity of ChatGPT and humans
par: Reviriego, Pedro, et autres
Publié: (2023)
par: Reviriego, Pedro, et autres
Publié: (2023)
Can large language models understand uncommon meanings of common words?
par: Wu, Jinyang, et autres
Publié: (2024)
par: Wu, Jinyang, et autres
Publié: (2024)
The Good, The Bad, and Why: Unveiling Emotions in Generative AI
par: Li, Cheng, et autres
Publié: (2023)
par: Li, Cheng, et autres
Publié: (2023)
Chain-of-Description: What I can understand, I can put into words
par: Guo, Jiaxin, et autres
Publié: (2025)
par: Guo, Jiaxin, et autres
Publié: (2025)
Scaling few-shot spoken word classification with generative meta-continual learning
par: Beyers, Louise, et autres
Publié: (2026)
par: Beyers, Louise, et autres
Publié: (2026)
How word semantics and phonology affect handwriting of Alzheimer's patients: a machine learning based analysis
par: Cilia, Nicole Dalia, et autres
Publié: (2023)
par: Cilia, Nicole Dalia, et autres
Publié: (2023)
Exploiting the English Vocabulary Profile for L2 word-level vocabulary assessment with LLMs
par: Bannò, Stefano, et autres
Publié: (2025)
par: Bannò, Stefano, et autres
Publié: (2025)
You shall know a piece by the company it keeps. Chess plays as a data for word2vec models
par: Orekhov, Boris
Publié: (2024)
par: Orekhov, Boris
Publié: (2024)
Why They Disagree: Decoding Differences in Opinions about AI Risk on the Lex Fridman Podcast
par: Truong, Nghi, et autres
Publié: (2025)
par: Truong, Nghi, et autres
Publié: (2025)
Choices Speak Louder than Questions
par: Cho, Gyeongje, et autres
Publié: (2025)
par: Cho, Gyeongje, et autres
Publié: (2025)
Does language matter for spoken word classification? A multilingual generative meta-learning approach
par: Ziki, Batsirayi Mupamhi, et autres
Publié: (2026)
par: Ziki, Batsirayi Mupamhi, et autres
Publié: (2026)
Why are LLMs' abilities emergent?
par: Havlík, Vladimír
Publié: (2025)
par: Havlík, Vladimír
Publié: (2025)
Plain language adaptations of biomedical text using LLMs: Comparision of evaluation metrics
par: Kocbek, Primoz, et autres
Publié: (2025)
par: Kocbek, Primoz, et autres
Publié: (2025)
Towards a resource for multilingual lexicons: an MT assisted and human-in-the-loop multilingual parallel corpus with multi-word expression annotation
par: Han, Lifeng, et autres
Publié: (2020)
par: Han, Lifeng, et autres
Publié: (2020)
MediFact at MEDIQA-CORR 2024: Why AI Needs a Human Touch
par: Saeed, Nadia
Publié: (2024)
par: Saeed, Nadia
Publié: (2024)
Why Attend to Everything? Focus is the Key
par: Yao, Hengshuai, et autres
Publié: (2026)
par: Yao, Hengshuai, et autres
Publié: (2026)
Meaningless is better: hashing bias-inducing words in LLM prompts improves performance in logical reasoning and statistical learning
par: Chadimová, Milena, et autres
Publié: (2024)
par: Chadimová, Milena, et autres
Publié: (2024)
Yes, this is what I was looking for! Towards Multi-modal Medical Consultation Concern Summary Generation
par: Tiwari, Abhisek, et autres
Publié: (2024)
par: Tiwari, Abhisek, et autres
Publié: (2024)
Why Slop Matters
par: Kommers, Cody, et autres
Publié: (2025)
par: Kommers, Cody, et autres
Publié: (2025)
LLM Olympiad: Why Model Evaluation Needs a Sealed Exam
par: Cruz, Jan Christian Blaise, et autres
Publié: (2026)
par: Cruz, Jan Christian Blaise, et autres
Publié: (2026)
Impact of enriched meaning representations for language generation in dialogue tasks: A comprehensive exploration of the relevance of tasks, corpora and metrics
par: Vázquez, Alain, et autres
Publié: (2026)
par: Vázquez, Alain, et autres
Publié: (2026)
Re-evaluating Theory of Mind evaluation in large language models
par: Hu, Jennifer, et autres
Publié: (2025)
par: Hu, Jennifer, et autres
Publié: (2025)
ks-lit-3m: A 3.1 million word kashmiri text dataset for large language model pretraining
par: Malik, Haq Nawaz
Publié: (2026)
par: Malik, Haq Nawaz
Publié: (2026)
From communities to interpretable network and word embedding: an unified approach
par: Prouteau, Thibault, et autres
Publié: (2024)
par: Prouteau, Thibault, et autres
Publié: (2024)
Development and multi-center evaluation of domain-adapted speech recognition for human-AI teaming in real-world gastrointestinal endoscopy
par: Yang, Ruijie, et autres
Publié: (2026)
par: Yang, Ruijie, et autres
Publié: (2026)
Why Chain of Thought Fails in Clinical Text Understanding
par: Wu, Jiageng, et autres
Publié: (2025)
par: Wu, Jiageng, et autres
Publié: (2025)
Documents similaires
-
ConSens: Assessing context grounding in open-book question answering
par: Vankov, Ivan, et autres
Publié: (2025) -
The Boy Who Survived: Removing Harry Potter from an LLM is harder than reported
par: Shostack, Adam
Publié: (2024) -
Evaluating Embedding Frameworks for Scientific Domain
par: Ahmed, Nouman, et autres
Publié: (2025) -
A word association network methodology for evaluating implicit biases in LLMs compared to humans
par: Abramski, Katherine, et autres
Publié: (2025) -
Automated alignment is harder than you think
par: Bowkis, Aleksandr, et autres
Publié: (2026)