Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Balamurali, Sai Shridhar, Cheng, Lu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering
by: Arias-Duart, Anna, et al.
Published: (2025)
by: Arias-Duart, Anna, et al.
Published: (2025)
Prompt-Based Clarity Evaluation and Topic Detection in Political Question Answering
by: Prahallad, Lavanya, et al.
Published: (2026)
by: Prahallad, Lavanya, et al.
Published: (2026)
RPO-RAG: Aligning Small LLMs with Relation-aware Preference Optimization for Knowledge Graph Question Answering
by: Um, Kaehyun, et al.
Published: (2026)
by: Um, Kaehyun, et al.
Published: (2026)
Aligning LLMs through Multi-perspective User Preference Ranking-based Feedback for Programming Question Answering
by: Yang, Hongyu, et al.
Published: (2024)
by: Yang, Hongyu, et al.
Published: (2024)
Knowledgeable Preference Alignment for LLMs in Domain-specific Question Answering
by: Zhang, Yichi, et al.
Published: (2023)
by: Zhang, Yichi, et al.
Published: (2023)
Evaluation of RAG Metrics for Question Answering in the Telecom Domain
by: Roychowdhury, Sujoy, et al.
Published: (2024)
by: Roychowdhury, Sujoy, et al.
Published: (2024)
QPaug: Question and Passage Augmentation for Open-Domain Question Answering of LLMs
by: Kim, Minsang, et al.
Published: (2024)
by: Kim, Minsang, et al.
Published: (2024)
SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
by: Kendre, Shrikant, et al.
Published: (2025)
by: Kendre, Shrikant, et al.
Published: (2025)
RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content
by: Monteiro, Joao, et al.
Published: (2024)
by: Monteiro, Joao, et al.
Published: (2024)
Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answering
by: Adlakha, Vaibhav, et al.
Published: (2023)
by: Adlakha, Vaibhav, et al.
Published: (2023)
Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics
by: Cho, Yousang, et al.
Published: (2025)
by: Cho, Yousang, et al.
Published: (2025)
Distilling LLMs' Decomposition Abilities into Compact Language Models
by: Tarasov, Denis, et al.
Published: (2024)
by: Tarasov, Denis, et al.
Published: (2024)
Cost-efficient Knowledge-based Question Answering with Large Language Models
by: Dong, Junnan, et al.
Published: (2024)
by: Dong, Junnan, et al.
Published: (2024)
MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
by: Wei, Jianhui, et al.
Published: (2025)
by: Wei, Jianhui, et al.
Published: (2025)
Affective-NLI: Towards Accurate and Interpretable Personality Recognition in Conversation
by: Wen, Zhiyuan, et al.
Published: (2024)
by: Wen, Zhiyuan, et al.
Published: (2024)
Fine-Tuning LLMs for Reliable Medical Question-Answering Services
by: Anaissi, Ali, et al.
Published: (2024)
by: Anaissi, Ali, et al.
Published: (2024)
HeySQuAD: A Spoken Question Answering Dataset
by: Wu, Yijing, et al.
Published: (2023)
by: Wu, Yijing, et al.
Published: (2023)
Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
by: Banayeeanzade, Amin, et al.
Published: (2025)
by: Banayeeanzade, Amin, et al.
Published: (2025)
Question Answering on Patient Medical Records with Private Fine-Tuned LLMs
by: Kothari, Sara, et al.
Published: (2025)
by: Kothari, Sara, et al.
Published: (2025)
Hybrid Graphs for Table-and-Text based Question Answering using LLMs
by: Agarwal, Ankush, et al.
Published: (2025)
by: Agarwal, Ankush, et al.
Published: (2025)
Augmenting Black-box LLMs with Medical Textbooks for Biomedical Question Answering
by: Wang, Yubo, et al.
Published: (2023)
by: Wang, Yubo, et al.
Published: (2023)
Training LLMs with Reinforcement Learning for Intent-Aware Personalized Question Answering
by: Amirizaniani, Maryam, et al.
Published: (2026)
by: Amirizaniani, Maryam, et al.
Published: (2026)
60 Data Points are Sufficient to Fine-Tune LLMs for Question-Answering
by: Ye, Junjie, et al.
Published: (2024)
by: Ye, Junjie, et al.
Published: (2024)
Performance Assessment of ChatGPT vs Bard in Detecting Alzheimer's Dementia
by: T, Balamurali B, et al.
Published: (2024)
by: T, Balamurali B, et al.
Published: (2024)
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation
by: Yang, Gao, et al.
Published: (2025)
by: Yang, Gao, et al.
Published: (2025)
Towards Human-Like Grading: A Unified LLM-Enhanced Framework for Subjective Question Evaluation
by: Zhua, Fanwei, et al.
Published: (2025)
by: Zhua, Fanwei, et al.
Published: (2025)
To Generate or to Retrieve? On the Effectiveness of Artificial Contexts for Medical Open-Domain Question Answering
by: Frisoni, Giacomo, et al.
Published: (2024)
by: Frisoni, Giacomo, et al.
Published: (2024)
Question Answering with LLMs and Learning from Answer Sets
by: Borroto, Manuel, et al.
Published: (2025)
by: Borroto, Manuel, et al.
Published: (2025)
Mind the Ambiguity: Aleatoric Uncertainty Quantification in LLMs for Safe Medical Question Answering
by: Liu, Yaokun, et al.
Published: (2026)
by: Liu, Yaokun, et al.
Published: (2026)
I've got the "Answer"! Interpretation of LLMs Hidden States in Question Answering
by: Goloviznina, Valeriya, et al.
Published: (2024)
by: Goloviznina, Valeriya, et al.
Published: (2024)
Generative Data Augmentation using LLMs improves Distributional Robustness in Question Answering
by: Chowdhury, Arijit Ghosh, et al.
Published: (2023)
by: Chowdhury, Arijit Ghosh, et al.
Published: (2023)
Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answering
by: Calvo-Bartolomé, Lorena, et al.
Published: (2025)
by: Calvo-Bartolomé, Lorena, et al.
Published: (2025)
Towards Robust Extractive Question Answering Models: Rethinking the Training Methodology
by: Tran, Son Quoc, et al.
Published: (2024)
by: Tran, Son Quoc, et al.
Published: (2024)
Towards Question Answering over Large Semi-structured Tables
by: Wang, Yuxiang, et al.
Published: (2025)
by: Wang, Yuxiang, et al.
Published: (2025)
TQA-Bench: Evaluating LLMs for Multi-Table Question Answering with Scalable Context and Symbolic Extension
by: Qiu, Zipeng, et al.
Published: (2024)
by: Qiu, Zipeng, et al.
Published: (2024)
A Learn-Then-Reason Model Towards Generalization in Knowledge Base Question Answering
by: Zhang, Lingxi, et al.
Published: (2024)
by: Zhang, Lingxi, et al.
Published: (2024)
Towards Better Generalization in Open-Domain Question Answering by Mitigating Context Memorization
by: Zhang, Zixuan, et al.
Published: (2024)
by: Zhang, Zixuan, et al.
Published: (2024)
Towards Robust Evaluation: A Comprehensive Taxonomy of Datasets and Metrics for Open Domain Question Answering in the Era of Large Language Models
by: Srivastava, Akchay, et al.
Published: (2024)
by: Srivastava, Akchay, et al.
Published: (2024)
PRISM: Agentic Retrieval with LLMs for Multi-Hop Question Answering
by: Nahid, Md Mahadi Hasan, et al.
Published: (2025)
by: Nahid, Md Mahadi Hasan, et al.
Published: (2025)
Graph Neural Network Enhanced Retrieval for Question Answering of LLMs
by: Li, Zijian, et al.
Published: (2024)
by: Li, Zijian, et al.
Published: (2024)
Similar Items
-
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering
by: Arias-Duart, Anna, et al.
Published: (2025) -
Prompt-Based Clarity Evaluation and Topic Detection in Political Question Answering
by: Prahallad, Lavanya, et al.
Published: (2026) -
RPO-RAG: Aligning Small LLMs with Relation-aware Preference Optimization for Knowledge Graph Question Answering
by: Um, Kaehyun, et al.
Published: (2026) -
Aligning LLMs through Multi-perspective User Preference Ranking-based Feedback for Programming Question Answering
by: Yang, Hongyu, et al.
Published: (2024) -
Knowledgeable Preference Alignment for LLMs in Domain-specific Question Answering
by: Zhang, Yichi, et al.
Published: (2023)