SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kendre, Shrikant, Xu, Austin, Zhou, Honglu, Ryoo, Michael, Joty, Shafiq, Niebles, Juan Carlos
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908669037248512
author Kendre, Shrikant
Xu, Austin
Zhou, Honglu
Ryoo, Michael
Joty, Shafiq
Niebles, Juan Carlos
author_facet Kendre, Shrikant
Xu, Austin
Zhou, Honglu
Ryoo, Michael
Joty, Shafiq
Niebles, Juan Carlos
contents Traditional evaluation metrics for textual and visual question answering, like ROUGE, METEOR, and Exact Match (EM), focus heavily on n-gram based lexical similarity, often missing the deeper semantic understanding needed for accurate assessment. While measures like BERTScore and MoverScore leverage contextual embeddings to address this limitation, they lack flexibility in balancing sentence-level and keyword-level semantics and ignore lexical similarity, which remains important. Large Language Model (LLM) based evaluators, though powerful, come with drawbacks like high costs, bias, inconsistency, and hallucinations. To address these issues, we introduce SMILE: Semantic Metric Integrating Lexical Exactness, a novel approach that combines sentence-level semantic understanding with keyword-level semantic understanding and easy keyword matching. This composite method balances lexical precision and semantic relevance, offering a comprehensive evaluation. Extensive benchmarks across text, image, and video QA tasks show SMILE is highly correlated with human judgments and computationally lightweight, bridging the gap between lexical and semantic evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2511_17432
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
Kendre, Shrikant
Xu, Austin
Zhou, Honglu
Ryoo, Michael
Joty, Shafiq
Niebles, Juan Carlos
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Traditional evaluation metrics for textual and visual question answering, like ROUGE, METEOR, and Exact Match (EM), focus heavily on n-gram based lexical similarity, often missing the deeper semantic understanding needed for accurate assessment. While measures like BERTScore and MoverScore leverage contextual embeddings to address this limitation, they lack flexibility in balancing sentence-level and keyword-level semantics and ignore lexical similarity, which remains important. Large Language Model (LLM) based evaluators, though powerful, come with drawbacks like high costs, bias, inconsistency, and hallucinations. To address these issues, we introduce SMILE: Semantic Metric Integrating Lexical Exactness, a novel approach that combines sentence-level semantic understanding with keyword-level semantic understanding and easy keyword matching. This composite method balances lexical precision and semantic relevance, offering a comprehensive evaluation. Extensive benchmarks across text, image, and video QA tasks show SMILE is highly correlated with human judgments and computationally lightweight, bridging the gap between lexical and semantic evaluation.
title SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.17432