Data-efficient Meta-models for Evaluation of Context-based Questions and Answers in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Belikova, Julia, Polev, Konstantin, Parchiev, Rauf, Simakov, Dmitry |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hallucination Detection in LLMs with Topological Divergence on Attention Graphs
von: Bazarova, Alexandra, et al.
Veröffentlicht: (2025)
von: Bazarova, Alexandra, et al.
Veröffentlicht: (2025)
Detecting Overflow in Compressed Token Representations for Retrieval-Augmented Generation
von: Belikova, Julia, et al.
Veröffentlicht: (2026)
von: Belikova, Julia, et al.
Veröffentlicht: (2026)
Probabilistic distances-based hallucination detection in LLMs with RAG
von: Oblovatny, Rodion, et al.
Veröffentlicht: (2025)
von: Oblovatny, Rodion, et al.
Veröffentlicht: (2025)
From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation
von: Zhou, Chengliang, et al.
Veröffentlicht: (2025)
von: Zhou, Chengliang, et al.
Veröffentlicht: (2025)
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
von: Veuthey, Jaime Raldua, et al.
Veröffentlicht: (2025)
von: Veuthey, Jaime Raldua, et al.
Veröffentlicht: (2025)
LLMs Provide Unstable Answers to Legal Questions
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2025)
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2025)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
Question Answering with LLMs and Learning from Answer Sets
von: Borroto, Manuel, et al.
Veröffentlicht: (2025)
von: Borroto, Manuel, et al.
Veröffentlicht: (2025)
Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers
von: Wang, Yuhan, et al.
Veröffentlicht: (2026)
von: Wang, Yuhan, et al.
Veröffentlicht: (2026)
Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering
von: Nachshoni, Eviatar, et al.
Veröffentlicht: (2025)
von: Nachshoni, Eviatar, et al.
Veröffentlicht: (2025)
The Potential of LLMs in Medical Education: Generating Questions and Answers for Qualification Exams
von: Zhu, Yunqi, et al.
Veröffentlicht: (2024)
von: Zhu, Yunqi, et al.
Veröffentlicht: (2024)
SAND-Math: Using LLMs to Generate Novel, Difficult and Useful Mathematics Questions and Answers
von: Manem, Chaitanya, et al.
Veröffentlicht: (2025)
von: Manem, Chaitanya, et al.
Veröffentlicht: (2025)
EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos
von: Ray, Sourjyadip, et al.
Veröffentlicht: (2025)
von: Ray, Sourjyadip, et al.
Veröffentlicht: (2025)
Evaluating Answer Reranking Strategies in Time-sensitive Question Answering
von: Kardan, Mehmet, et al.
Veröffentlicht: (2025)
von: Kardan, Mehmet, et al.
Veröffentlicht: (2025)
Ask, Answer, and Detect: Role-Playing LLMs for Personality Detection with Question-Conditioned Mixture-of-Experts
von: Lyu, Yifan, et al.
Veröffentlicht: (2025)
von: Lyu, Yifan, et al.
Veröffentlicht: (2025)
I've got the "Answer"! Interpretation of LLMs Hidden States in Question Answering
von: Goloviznina, Valeriya, et al.
Veröffentlicht: (2024)
von: Goloviznina, Valeriya, et al.
Veröffentlicht: (2024)
Putting People in LLMs' Shoes: Generating Better Answers via Question Rewriter
von: Chen, Junhao, et al.
Veröffentlicht: (2024)
von: Chen, Junhao, et al.
Veröffentlicht: (2024)
Automatic Feedback Generation for Short Answer Questions using Answer Diagnostic Graphs
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2025)
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2025)
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
von: Naeem, Numaan, et al.
Veröffentlicht: (2025)
von: Naeem, Numaan, et al.
Veröffentlicht: (2025)
Sparse Autoencoders for Sequential Recommendation Models: Interpretation and Flexible Control
von: Klenitskiy, Anton, et al.
Veröffentlicht: (2025)
von: Klenitskiy, Anton, et al.
Veröffentlicht: (2025)
Knowledge-Augmented Question Error Correction for Chinese Question Answer System with QuestionRAG
von: Qiu, Longpeng, et al.
Veröffentlicht: (2025)
von: Qiu, Longpeng, et al.
Veröffentlicht: (2025)
Evaluating Robustness of LLMs in Question Answering on Multilingual Noisy OCR Data
von: Piryani, Bhawna, et al.
Veröffentlicht: (2025)
von: Piryani, Bhawna, et al.
Veröffentlicht: (2025)
Rehearsing Answers to Probable Questions with Perspective-Taking
von: Shih, Yung-Yu, et al.
Veröffentlicht: (2024)
von: Shih, Yung-Yu, et al.
Veröffentlicht: (2024)
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
von: Yona, Gal, et al.
Veröffentlicht: (2024)
von: Yona, Gal, et al.
Veröffentlicht: (2024)
QGen Studio: An Adaptive Question-Answer Generation, Training and Evaluation Platform
von: Moses, Movina, et al.
Veröffentlicht: (2025)
von: Moses, Movina, et al.
Veröffentlicht: (2025)
Evaluation Methodology for Large Language Models for Multilingual Document Question and Answer
von: Kahana, Adar, et al.
Veröffentlicht: (2024)
von: Kahana, Adar, et al.
Veröffentlicht: (2024)
Comparative Analysis of 47 Context-Based Question Answer Models Across 8 Diverse Datasets
von: Muneeb, Muhammad, et al.
Veröffentlicht: (2025)
von: Muneeb, Muhammad, et al.
Veröffentlicht: (2025)
Evaluating Biases in Context-Dependent Health Questions
von: Levy, Sharon, et al.
Veröffentlicht: (2024)
von: Levy, Sharon, et al.
Veröffentlicht: (2024)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2025)
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2025)
A Study on Question-Answer Dataset for LLM Safety Evaluation with a Focus on Illegal Activities
von: Imamura, Kenji, et al.
Veröffentlicht: (2026)
von: Imamura, Kenji, et al.
Veröffentlicht: (2026)
CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering
von: Li, Zongxia, et al.
Veröffentlicht: (2024)
von: Li, Zongxia, et al.
Veröffentlicht: (2024)
GETALP@AutoMin 2025: Leveraging RAG to Answer Questions based on Meeting Transcripts
von: Kang, Jeongwoo, et al.
Veröffentlicht: (2025)
von: Kang, Jeongwoo, et al.
Veröffentlicht: (2025)
Do LLMs Provide Consistent Answers to Health-Related Questions across Languages?
von: Schlicht, Ipek Baris, et al.
Veröffentlicht: (2025)
von: Schlicht, Ipek Baris, et al.
Veröffentlicht: (2025)
Controllable Decontextualization of Yes/No Question and Answers into Factual Statements
von: Mo, Lingbo, et al.
Veröffentlicht: (2024)
von: Mo, Lingbo, et al.
Veröffentlicht: (2024)
Automatic Question-Answer Generation for Long-Tail Knowledge
von: Kumar, Rohan, et al.
Veröffentlicht: (2024)
von: Kumar, Rohan, et al.
Veröffentlicht: (2024)
TARAZ: Persian Short-Answer Question Benchmark for Cultural Evaluation of Language Models
von: Iranmanesh, Reihaneh, et al.
Veröffentlicht: (2026)
von: Iranmanesh, Reihaneh, et al.
Veröffentlicht: (2026)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
von: Wiegreffe, Sarah, et al.
Veröffentlicht: (2024)
von: Wiegreffe, Sarah, et al.
Veröffentlicht: (2024)
When Answers Stray from Questions: Hallucination Detection via Question-Answer Orthogonal Decomposition
von: Yao, Siyang, et al.
Veröffentlicht: (2026)
von: Yao, Siyang, et al.
Veröffentlicht: (2026)
Question: How do Large Language Models perform on the Question Answering tasks? Answer:
von: Fischer, Kevin, et al.
Veröffentlicht: (2024)
von: Fischer, Kevin, et al.
Veröffentlicht: (2024)
BPQA Dataset: Evaluating How Well Language Models Leverage Blood Pressures to Answer Biomedical Questions
von: Hang, Chi, et al.
Veröffentlicht: (2025)
von: Hang, Chi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hallucination Detection in LLMs with Topological Divergence on Attention Graphs
von: Bazarova, Alexandra, et al.
Veröffentlicht: (2025) -
Detecting Overflow in Compressed Token Representations for Retrieval-Augmented Generation
von: Belikova, Julia, et al.
Veröffentlicht: (2026) -
Probabilistic distances-based hallucination detection in LLMs with RAG
von: Oblovatny, Rodion, et al.
Veröffentlicht: (2025) -
From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation
von: Zhou, Chengliang, et al.
Veröffentlicht: (2025) -
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
von: Veuthey, Jaime Raldua, et al.
Veröffentlicht: (2025)