Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Adlakha, Vaibhav, BehnamGhader, Parishad, Lu, Xing Han, Meade, Nicholas, Reddy, Siva |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploiting Instruction-Following Retrievers for Malicious Information Retrieval
by: BehnamGhader, Parishad, et al.
Published: (2025)
by: BehnamGhader, Parishad, et al.
Published: (2025)
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
by: BehnamGhader, Parishad, et al.
Published: (2024)
by: BehnamGhader, Parishad, et al.
Published: (2024)
LLM2Vec-Gen: Generative Embeddings from Large Language Models
by: BehnamGhader, Parishad, et al.
Published: (2026)
by: BehnamGhader, Parishad, et al.
Published: (2026)
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
by: Marjanović, Sara Vera, et al.
Published: (2025)
by: Marjanović, Sara Vera, et al.
Published: (2025)
SafeArena: Evaluating the Safety of Autonomous Web Agents
by: Tur, Ada Defne, et al.
Published: (2025)
by: Tur, Ada Defne, et al.
Published: (2025)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
by: Lù, Xing Han, et al.
Published: (2025)
by: Lù, Xing Han, et al.
Published: (2025)
Multi-hop Question Answering
by: Mavi, Vaibhav, et al.
Published: (2022)
by: Mavi, Vaibhav, et al.
Published: (2022)
From RAG to Agentic RAG for Faithful Islamic Question Answering
by: Bhatia, Gagan, et al.
Published: (2026)
by: Bhatia, Gagan, et al.
Published: (2026)
FiDeLiS: Faithful Reasoning in Large Language Model for Knowledge Graph Question Answering
by: Sui, Yuan, et al.
Published: (2024)
by: Sui, Yuan, et al.
Published: (2024)
Towards Faithful Knowledge Graph Explanation Through Deep Alignment in Commonsense Question Answering
by: Zhai, Weihe, et al.
Published: (2023)
by: Zhai, Weihe, et al.
Published: (2023)
An Analysis of Decoding Methods for LLM-based Agents for Faithful Multi-Hop Question Answering
by: Murphy, Alexander, et al.
Published: (2025)
by: Murphy, Alexander, et al.
Published: (2025)
Investigating Adversarial Trigger Transfer in Large Language Models
by: Meade, Nicholas, et al.
Published: (2024)
by: Meade, Nicholas, et al.
Published: (2024)
Self-Correction Distillation for Structured Data Question Answering
by: Zhu, Yushan, et al.
Published: (2025)
by: Zhu, Yushan, et al.
Published: (2025)
Optimizing Medical Question-Answering Systems: A Comparative Study of Fine-Tuned and Zero-Shot Large Language Models with RAG Framework
by: Hassan, Tasnimul, et al.
Published: (2025)
by: Hassan, Tasnimul, et al.
Published: (2025)
Exploring the Robustness of Language Models for Tabular Question Answering via Attention Analysis
by: Bhandari, Kushal Raj, et al.
Published: (2024)
by: Bhandari, Kushal Raj, et al.
Published: (2024)
Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models
by: Sun, Wangtao, et al.
Published: (2024)
by: Sun, Wangtao, et al.
Published: (2024)
LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
by: Ren, Huimin, et al.
Published: (2025)
by: Ren, Huimin, et al.
Published: (2025)
Mitigating Knowledge Conflicts in Language Model-Driven Question Answering
by: Cao, Han, et al.
Published: (2024)
by: Cao, Han, et al.
Published: (2024)
Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering
by: Balamurali, Sai Shridhar, et al.
Published: (2025)
by: Balamurali, Sai Shridhar, et al.
Published: (2025)
Omne-R1: Learning to Reason with Memory for Multi-hop Question Answering
by: Liu, Boyuan, et al.
Published: (2025)
by: Liu, Boyuan, et al.
Published: (2025)
Towards Better Generalization in Open-Domain Question Answering by Mitigating Context Memorization
by: Zhang, Zixuan, et al.
Published: (2024)
by: Zhang, Zixuan, et al.
Published: (2024)
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering
by: Arias-Duart, Anna, et al.
Published: (2025)
by: Arias-Duart, Anna, et al.
Published: (2025)
Evaluating the Meta- and Object-Level Reasoning of Large Language Models for Question Answering
by: Ferguson, Nick, et al.
Published: (2025)
by: Ferguson, Nick, et al.
Published: (2025)
LIFEBench: Evaluating Length Instruction Following in Large Language Models
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
by: Krojer, Benno, et al.
Published: (2026)
by: Krojer, Benno, et al.
Published: (2026)
Language Models Largely Exhibit Human-like Constituent Ordering Preferences
by: Tur, Ada Defne, et al.
Published: (2025)
by: Tur, Ada Defne, et al.
Published: (2025)
GSQA: An End-to-End Model for Generative Spoken Question Answering
by: Shih, Min-Han, et al.
Published: (2023)
by: Shih, Min-Han, et al.
Published: (2023)
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
ORCA: Open-ended Response Correctness Assessment for Audio Question Answering
by: Sedláček, Šimon, et al.
Published: (2025)
by: Sedláček, Šimon, et al.
Published: (2025)
Understanding the Influence of Synthetic Data for Text Embedders
by: Springer, Jacob Mitchell, et al.
Published: (2025)
by: Springer, Jacob Mitchell, et al.
Published: (2025)
M-IFEval: Multilingual Instruction-Following Evaluation
by: Dussolle, Antoine, et al.
Published: (2025)
by: Dussolle, Antoine, et al.
Published: (2025)
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models
by: Lyu, Xinxi, et al.
Published: (2024)
by: Lyu, Xinxi, et al.
Published: (2024)
RefuteBench: Evaluating Refuting Instruction-Following for Large Language Models
by: Yan, Jianhao, et al.
Published: (2024)
by: Yan, Jianhao, et al.
Published: (2024)
Scope Ambiguities in Large Language Models
by: Kamath, Gaurav, et al.
Published: (2024)
by: Kamath, Gaurav, et al.
Published: (2024)
Condition-Gated Reasoning for Context-Dependent Biomedical Question Answering
by: Parekh, Jash Rajesh, et al.
Published: (2026)
by: Parekh, Jash Rajesh, et al.
Published: (2026)
Training LLMs with Reinforcement Learning for Intent-Aware Personalized Question Answering
by: Amirizaniani, Maryam, et al.
Published: (2026)
by: Amirizaniani, Maryam, et al.
Published: (2026)
Understanding Network Behaviors through Natural Language Question-Answering
by: Xing, Mingzhe, et al.
Published: (2025)
by: Xing, Mingzhe, et al.
Published: (2025)
Correct after Answer: Enhancing Multi-Span Question Answering with Post-Processing Method
by: Lin, Jiayi, et al.
Published: (2024)
by: Lin, Jiayi, et al.
Published: (2024)
TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering
by: Han, Liying, et al.
Published: (2026)
by: Han, Liying, et al.
Published: (2026)
Are self-explanations from Large Language Models faithful?
by: Madsen, Andreas, et al.
Published: (2024)
by: Madsen, Andreas, et al.
Published: (2024)
Similar Items
-
Exploiting Instruction-Following Retrievers for Malicious Information Retrieval
by: BehnamGhader, Parishad, et al.
Published: (2025) -
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
by: BehnamGhader, Parishad, et al.
Published: (2024) -
LLM2Vec-Gen: Generative Embeddings from Large Language Models
by: BehnamGhader, Parishad, et al.
Published: (2026) -
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
by: Marjanović, Sara Vera, et al.
Published: (2025) -
SafeArena: Evaluating the Safety of Autonomous Web Agents
by: Tur, Ada Defne, et al.
Published: (2025)