Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses
Fuente:
arXiv
Saved in:
| Main Authors: | Ho, Xanh, Huang, Jiahao, Boudin, Florian, Aizawa, Akiko |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoreHopQA: More Than Multi-hop Reasoning
by: Schnitzler, Julian, et al.
Published: (2024)
by: Schnitzler, Julian, et al.
Published: (2024)
An Analysis of Datasets, Metrics and Models in Keyphrase Generation
by: Boudin, Florian, et al.
Published: (2025)
by: Boudin, Florian, et al.
Published: (2025)
Unsupervised Domain Adaptation for Keyphrase Generation using Citation Contexts
by: Boudin, Florian, et al.
Published: (2024)
by: Boudin, Florian, et al.
Published: (2024)
Table-Text Alignment: Explaining Claim Verification Against Tables in Scientific Papers
by: Ho, Xanh, et al.
Published: (2025)
by: Ho, Xanh, et al.
Published: (2025)
Format Matters: The Robustness of Multimodal LLMs in Reviewing Evidence from Tables and Charts
by: Ho, Xanh, et al.
Published: (2025)
by: Ho, Xanh, et al.
Published: (2025)
Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification
by: Kumar, Sunisth, et al.
Published: (2026)
by: Kumar, Sunisth, et al.
Published: (2026)
Preface to the Special Issue of the TAL Journal on Scholarly Document Processing
by: Boudin, Florian, et al.
Published: (2025)
by: Boudin, Florian, et al.
Published: (2025)
Automatically Suggesting Diverse Example Sentences for L2 Japanese Learners Using Pre-Trained Language Models
by: Benedetti, Enrico, et al.
Published: (2025)
by: Benedetti, Enrico, et al.
Published: (2025)
SciClaimEval: Cross-modal Claim Verification in Scientific Papers
by: Ho, Xanh, et al.
Published: (2026)
by: Ho, Xanh, et al.
Published: (2026)
A Survey of Pre-trained Language Models for Processing Scientific Text
by: Ho, Xanh, et al.
Published: (2024)
by: Ho, Xanh, et al.
Published: (2024)
Self-Compositional Data Augmentation for Scientific Keyphrase Generation
by: Houbre, Mael, et al.
Published: (2024)
by: Houbre, Mael, et al.
Published: (2024)
ParaRev: Building a dataset for Scientific Paragraph Revision annotated with revision instruction
by: Jourdan, Léane, et al.
Published: (2025)
by: Jourdan, Léane, et al.
Published: (2025)
Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
by: Kim, Kon Woo, et al.
Published: (2025)
by: Kim, Kon Woo, et al.
Published: (2025)
JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models
by: Jiang, Junfeng, et al.
Published: (2024)
by: Jiang, Junfeng, et al.
Published: (2024)
FC-CONAN: An Exhaustively Paired Dataset for Robust Evaluation of Retrieval Systems
by: Junqueras, Juan, et al.
Published: (2026)
by: Junqueras, Juan, et al.
Published: (2026)
SKT5SciSumm -- Revisiting Extractive-Generative Approach for Multi-Document Scientific Summarization
by: To, Huy Quoc, et al.
Published: (2024)
by: To, Huy Quoc, et al.
Published: (2024)
Are Emotions Arranged in a Circle? Geometric Analysis of Emotion Representations via Hyperspherical Contrastive Learning
by: Yamauchi, Yusuke, et al.
Published: (2026)
by: Yamauchi, Yusuke, et al.
Published: (2026)
OMoS-QA: A Dataset for Cross-Lingual Extractive Question Answering in a German Migration Context
by: Kleinle, Steffen, et al.
Published: (2024)
by: Kleinle, Steffen, et al.
Published: (2024)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
by: Belmadani, Ikram, et al.
Published: (2026)
by: Belmadani, Ikram, et al.
Published: (2026)
Refining and Reusing Annotation Guidelines for LLM Annotation
by: Kim, Kon Woo, et al.
Published: (2026)
by: Kim, Kon Woo, et al.
Published: (2026)
Exploring Language Model Generalization in Low-Resource Extractive QA
by: Sengupta, Saptarshi, et al.
Published: (2024)
by: Sengupta, Saptarshi, et al.
Published: (2024)
Harnessing PDF Data for Improving Japanese Large Multimodal Models
by: Baek, Jeonghun, et al.
Published: (2025)
by: Baek, Jeonghun, et al.
Published: (2025)
The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias Detection
by: Horych, Tomas, et al.
Published: (2024)
by: Horych, Tomas, et al.
Published: (2024)
PrionNER: A Named Entity Recognition Dataset for Prion Disease Biomedical Literature
by: Dao, An, et al.
Published: (2026)
by: Dao, An, et al.
Published: (2026)
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
by: Lee, Dongryeol, et al.
Published: (2026)
by: Lee, Dongryeol, et al.
Published: (2026)
From Multiple-Choice to Extractive QA: A Case Study for English and Arabic
by: Lynn, Teresa, et al.
Published: (2024)
by: Lynn, Teresa, et al.
Published: (2024)
Evaluating the Homogeneity of Keyphrase Prediction Models
by: Houbre, Maël, et al.
Published: (2026)
by: Houbre, Maël, et al.
Published: (2026)
Through the LLM Looking Glass: A Socratic Probing of Donkeys, Elephants, and Markets
by: Kennedy, Molly, et al.
Published: (2025)
by: Kennedy, Molly, et al.
Published: (2025)
Few-Shot Prompting for Extractive Quranic QA with Instruction-Tuned LLMs
by: Basem, Mohamed, et al.
Published: (2025)
by: Basem, Mohamed, et al.
Published: (2025)
EarlySciRev: A Dataset of Early-Stage Scientific Revisions Extracted from LaTeX Writing Traces
by: Jourdan, Léane, et al.
Published: (2026)
by: Jourdan, Léane, et al.
Published: (2026)
ACL-rlg: A Dataset for Reading List Generation
by: Aubert-Béduchaud, Julien, et al.
Published: (2024)
by: Aubert-Béduchaud, Julien, et al.
Published: (2024)
Assessing Large Language Models for Medical QA: Zero-Shot and LLM-as-a-Judge Evaluation
by: Adib, Shefayat E Shams, et al.
Published: (2026)
by: Adib, Shefayat E Shams, et al.
Published: (2026)
Identifying Reliable Evaluation Metrics for Scientific Text Revision
by: Jourdan, Léane, et al.
Published: (2025)
by: Jourdan, Léane, et al.
Published: (2025)
CASIMIR: A Corpus of Scientific Articles enhanced with Multiple Author-Integrated Revisions
by: Jourdan, Leane, et al.
Published: (2024)
by: Jourdan, Leane, et al.
Published: (2024)
Text revision in Scientific Writing Assistance: An Overview
by: Jourdan, Léane, et al.
Published: (2023)
by: Jourdan, Léane, et al.
Published: (2023)
RJUA-QA: A Comprehensive QA Dataset for Urology
by: Lyu, Shiwei, et al.
Published: (2023)
by: Lyu, Shiwei, et al.
Published: (2023)
Tracing Multilingual Knowledge Acquisition Dynamics in Domain Adaptation: A Case Study of English-Japanese Biomedical Adaptation
by: Zhao, Xin, et al.
Published: (2025)
by: Zhao, Xin, et al.
Published: (2025)
Optimal Query Allocation in Extractive QA with LLMs: A Learning-to-Defer Framework with Theoretical Guarantees
by: Montreuil, Yannis, et al.
Published: (2024)
by: Montreuil, Yannis, et al.
Published: (2024)
Reasoning Depth and Environment Complexity: A Controlled Study of RLVR Data Allocation across Logical Reasoning Tasks
by: Zhu, Yihua, et al.
Published: (2026)
by: Zhu, Yihua, et al.
Published: (2026)
Rethinking Atomic Decomposition for LLM Judges: A Prompt-Controlled Study of Reference-Grounded QA Evaluation
by: Zhang, Xinran
Published: (2026)
by: Zhang, Xinran
Published: (2026)
Similar Items
-
MoreHopQA: More Than Multi-hop Reasoning
by: Schnitzler, Julian, et al.
Published: (2024) -
An Analysis of Datasets, Metrics and Models in Keyphrase Generation
by: Boudin, Florian, et al.
Published: (2025) -
Unsupervised Domain Adaptation for Keyphrase Generation using Citation Contexts
by: Boudin, Florian, et al.
Published: (2024) -
Table-Text Alignment: Explaining Claim Verification Against Tables in Scientific Papers
by: Ho, Xanh, et al.
Published: (2025) -
Format Matters: The Robustness of Multimodal LLMs in Reviewing Evidence from Tables and Charts
by: Ho, Xanh, et al.
Published: (2025)