Guardado en:
| Autores principales: | Imamura, Kenji, Ideuchi, Masao, Fujita, Atsushi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2605.29340 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CADEL: A Corpus of Administrative Web Documents for Japanese Entity Linking
por: Higashiyama, Shohei, et al.
Publicado: (2026)
por: Higashiyama, Shohei, et al.
Publicado: (2026)
ATD-Trans: A Geographically Grounded Japanese-English Travelogue Translation Dataset
por: Higashiyama, Shohei, et al.
Publicado: (2026)
por: Higashiyama, Shohei, et al.
Publicado: (2026)
AnswerCarefully: A Dataset for Improving the Safety of Japanese LLM Output
por: Suzuki, Hisami, et al.
Publicado: (2025)
por: Suzuki, Hisami, et al.
Publicado: (2025)
Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers
por: Wang, Yuhan, et al.
Publicado: (2026)
por: Wang, Yuhan, et al.
Publicado: (2026)
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
por: Veuthey, Jaime Raldua, et al.
Publicado: (2025)
por: Veuthey, Jaime Raldua, et al.
Publicado: (2025)
Focusing on Students, not Machines: Grounded Question Generation and Automated Answer Grading
por: Meyer, Gérôme, et al.
Publicado: (2025)
por: Meyer, Gérôme, et al.
Publicado: (2025)
Language-free Experience at Expo 2025 Osaka
por: Paul, Michael, et al.
Publicado: (2026)
por: Paul, Michael, et al.
Publicado: (2026)
BPQA Dataset: Evaluating How Well Language Models Leverage Blood Pressures to Answer Biomedical Questions
por: Hang, Chi, et al.
Publicado: (2025)
por: Hang, Chi, et al.
Publicado: (2025)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
por: Molfese, Francesco Maria, et al.
Publicado: (2025)
por: Molfese, Francesco Maria, et al.
Publicado: (2025)
A Dataset of Open-Domain Question Answering with Multiple-Span Answers
por: Luo, Zhiyi, et al.
Publicado: (2024)
por: Luo, Zhiyi, et al.
Publicado: (2024)
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
por: Cencerrado, Iván Vicente Moreno, et al.
Publicado: (2025)
por: Cencerrado, Iván Vicente Moreno, et al.
Publicado: (2025)
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
por: Balepur, Nishant, et al.
Publicado: (2024)
por: Balepur, Nishant, et al.
Publicado: (2024)
From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation
por: Zhou, Chengliang, et al.
Publicado: (2025)
por: Zhou, Chengliang, et al.
Publicado: (2025)
Can LLMs Grade Short-Answer Reading Comprehension Questions : An Empirical Study with a Novel Dataset
por: Henkel, Owen, et al.
Publicado: (2023)
por: Henkel, Owen, et al.
Publicado: (2023)
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
por: Naeem, Numaan, et al.
Publicado: (2025)
por: Naeem, Numaan, et al.
Publicado: (2025)
The Uneven Impact of Post-Training Quantization in Machine Translation
por: Marie, Benjamin, et al.
Publicado: (2025)
por: Marie, Benjamin, et al.
Publicado: (2025)
MiRAGE: A Multiagent Framework for Generating Multimodal Multihop Question-Answer Dataset for RAG Evaluation
por: Sahu, Chandan Kumar, et al.
Publicado: (2026)
por: Sahu, Chandan Kumar, et al.
Publicado: (2026)
Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering
por: Nachshoni, Eviatar, et al.
Publicado: (2025)
por: Nachshoni, Eviatar, et al.
Publicado: (2025)
Polyglots or Multitudes? Multilingual LLM Answers to Value-laden Multiple-Choice Questions
por: Labat, Léo, et al.
Publicado: (2026)
por: Labat, Léo, et al.
Publicado: (2026)
Evaluating Answer Reranking Strategies in Time-sensitive Question Answering
por: Kardan, Mehmet, et al.
Publicado: (2025)
por: Kardan, Mehmet, et al.
Publicado: (2025)
DEEPAMBIGQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness
por: Ji, Jiabao, et al.
Publicado: (2025)
por: Ji, Jiabao, et al.
Publicado: (2025)
Data-efficient Meta-models for Evaluation of Context-based Questions and Answers in LLMs
por: Belikova, Julia, et al.
Publicado: (2025)
por: Belikova, Julia, et al.
Publicado: (2025)
Automatic Feedback Generation for Short Answer Questions using Answer Diagnostic Graphs
por: Furuhashi, Momoka, et al.
Publicado: (2025)
por: Furuhashi, Momoka, et al.
Publicado: (2025)
Evaluation Methodology for Large Language Models for Multilingual Document Question and Answer
por: Kahana, Adar, et al.
Publicado: (2024)
por: Kahana, Adar, et al.
Publicado: (2024)
QGen Studio: An Adaptive Question-Answer Generation, Training and Evaluation Platform
por: Moses, Movina, et al.
Publicado: (2025)
por: Moses, Movina, et al.
Publicado: (2025)
Integrated Framework for LLM Evaluation with Answer Generation
por: Lee, Sujeong, et al.
Publicado: (2025)
por: Lee, Sujeong, et al.
Publicado: (2025)
Automatic Question & Answer Generation Using Generative Large Language Model (LLM)
por: Ehsan, Md. Alvee, et al.
Publicado: (2025)
por: Ehsan, Md. Alvee, et al.
Publicado: (2025)
Knowledge-Augmented Question Error Correction for Chinese Question Answer System with QuestionRAG
por: Qiu, Longpeng, et al.
Publicado: (2025)
por: Qiu, Longpeng, et al.
Publicado: (2025)
LLMs Provide Unstable Answers to Legal Questions
por: Blair-Stanek, Andrew, et al.
Publicado: (2025)
por: Blair-Stanek, Andrew, et al.
Publicado: (2025)
Rehearsing Answers to Probable Questions with Perspective-Taking
por: Shih, Yung-Yu, et al.
Publicado: (2024)
por: Shih, Yung-Yu, et al.
Publicado: (2024)
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
por: Yona, Gal, et al.
Publicado: (2024)
por: Yona, Gal, et al.
Publicado: (2024)
A Dataset for Evaluating LLM-based Evaluation Functions for Research Question Extraction Task
por: Fujisaki, Yuya, et al.
Publicado: (2024)
por: Fujisaki, Yuya, et al.
Publicado: (2024)
When Answers Stray from Questions: Hallucination Detection via Question-Answer Orthogonal Decomposition
por: Yao, Siyang, et al.
Publicado: (2026)
por: Yao, Siyang, et al.
Publicado: (2026)
Comparative Analysis of 47 Context-Based Question Answer Models Across 8 Diverse Datasets
por: Muneeb, Muhammad, et al.
Publicado: (2025)
por: Muneeb, Muhammad, et al.
Publicado: (2025)
TARAZ: Persian Short-Answer Question Benchmark for Cultural Evaluation of Language Models
por: Iranmanesh, Reihaneh, et al.
Publicado: (2026)
por: Iranmanesh, Reihaneh, et al.
Publicado: (2026)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
por: Wiegreffe, Sarah, et al.
Publicado: (2024)
por: Wiegreffe, Sarah, et al.
Publicado: (2024)
Decomposed Prompting to Answer Questions on a Course Discussion Board
por: Jaipersaud, Brandon, et al.
Publicado: (2024)
por: Jaipersaud, Brandon, et al.
Publicado: (2024)
Controllable Decontextualization of Yes/No Question and Answers into Factual Statements
por: Mo, Lingbo, et al.
Publicado: (2024)
por: Mo, Lingbo, et al.
Publicado: (2024)
Automatic Question-Answer Generation for Long-Tail Knowledge
por: Kumar, Rohan, et al.
Publicado: (2024)
por: Kumar, Rohan, et al.
Publicado: (2024)
CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering
por: Li, Zongxia, et al.
Publicado: (2024)
por: Li, Zongxia, et al.
Publicado: (2024)
Ejemplares similares
-
CADEL: A Corpus of Administrative Web Documents for Japanese Entity Linking
por: Higashiyama, Shohei, et al.
Publicado: (2026) -
ATD-Trans: A Geographically Grounded Japanese-English Travelogue Translation Dataset
por: Higashiyama, Shohei, et al.
Publicado: (2026) -
AnswerCarefully: A Dataset for Improving the Safety of Japanese LLM Output
por: Suzuki, Hisami, et al.
Publicado: (2025) -
Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers
por: Wang, Yuhan, et al.
Publicado: (2026) -
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
por: Veuthey, Jaime Raldua, et al.
Publicado: (2025)