BookAsSumQA: An Evaluation Framework for Aspect-Based Book Summarization via Question Answering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Miyazato, Ryuhei, Wei, Ting-Ruen, Wu, Xuyang, Wu, Hsin-Tai, Harada, Kei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EnsemHalDet: Robust VLM Hallucination Detection via Ensemble of Internal State Detectors
von: Miyazato, Ryuhei, et al.
Veröffentlicht: (2026)
von: Miyazato, Ryuhei, et al.
Veröffentlicht: (2026)
Does Reasoning Introduce Bias? A Study of Social Bias Evaluation and Mitigation in LLM Reasoning
von: Wu, Xuyang, et al.
Veröffentlicht: (2025)
von: Wu, Xuyang, et al.
Veröffentlicht: (2025)
Passage-specific Prompt Tuning for Passage Reranking in Question Answering with Large Language Models
von: Wu, Xuyang, et al.
Veröffentlicht: (2024)
von: Wu, Xuyang, et al.
Veröffentlicht: (2024)
Table Transformers for Imputing Textual Attributes
von: Wei, Ting-Ruen, et al.
Veröffentlicht: (2024)
von: Wei, Ting-Ruen, et al.
Veröffentlicht: (2024)
A Survey on Feedback-based Multi-step Reasoning for Large Language Models on Mathematics
von: Wei, Ting-Ruen, et al.
Veröffentlicht: (2025)
von: Wei, Ting-Ruen, et al.
Veröffentlicht: (2025)
DebateQA: Evaluating Question Answering on Debatable Knowledge
von: Xu, Rongwu, et al.
Veröffentlicht: (2024)
von: Xu, Rongwu, et al.
Veröffentlicht: (2024)
QA-prompting: Improving Summarization with Large Language Models using Question-Answering
von: Sinha, Neelabh
Veröffentlicht: (2025)
von: Sinha, Neelabh
Veröffentlicht: (2025)
Full-range Head Pose Geometric Data Augmentations
von: Hu, Huei-Chung, et al.
Veröffentlicht: (2024)
von: Hu, Huei-Chung, et al.
Veröffentlicht: (2024)
Does RAG Introduce Unfairness in LLMs? Evaluating Fairness in Retrieval-Augmented Generation Systems
von: Wu, Xuyang, et al.
Veröffentlicht: (2024)
von: Wu, Xuyang, et al.
Veröffentlicht: (2024)
MedExQA: Medical Question Answering Benchmark with Multiple Explanations
von: Kim, Yunsoo, et al.
Veröffentlicht: (2024)
von: Kim, Yunsoo, et al.
Veröffentlicht: (2024)
Evaluating Fairness in Large Vision-Language Models Across Diverse Demographic Attributes and Prompts
von: Wu, Xuyang, et al.
Veröffentlicht: (2024)
von: Wu, Xuyang, et al.
Veröffentlicht: (2024)
MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization
von: Nguyen, Huyen, et al.
Veröffentlicht: (2026)
von: Nguyen, Huyen, et al.
Veröffentlicht: (2026)
PolQA: Polish Question Answering Dataset
von: Rybak, Piotr, et al.
Veröffentlicht: (2022)
von: Rybak, Piotr, et al.
Veröffentlicht: (2022)
DisastQA: A Comprehensive Benchmark for Evaluating Question Answering in Disaster Management
von: Chen, Zhitong, et al.
Veröffentlicht: (2026)
von: Chen, Zhitong, et al.
Veröffentlicht: (2026)
CLERF: Contrastive LEaRning for Full Range Head Pose Estimation
von: Wei, Ting-Ruen, et al.
Veröffentlicht: (2024)
von: Wei, Ting-Ruen, et al.
Veröffentlicht: (2024)
ASTRA-QA: A Benchmark for Abstract Question Answering over Documents
von: Wang, Shu, et al.
Veröffentlicht: (2026)
von: Wang, Shu, et al.
Veröffentlicht: (2026)
Long-Span Question-Answering: Automatic Question Generation and QA-System Ranking via Side-by-Side Evaluation
von: Bohnet, Bernd, et al.
Veröffentlicht: (2024)
von: Bohnet, Bernd, et al.
Veröffentlicht: (2024)
Aspect-Based Summarization with Self-Aspect Retrieval Enhanced Generation
von: Feng, Yichao, et al.
Veröffentlicht: (2025)
von: Feng, Yichao, et al.
Veröffentlicht: (2025)
MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering
von: Islamaj, Rezarta, et al.
Veröffentlicht: (2026)
von: Islamaj, Rezarta, et al.
Veröffentlicht: (2026)
UserSumBench: A Benchmark Framework for Evaluating User Summarization Approaches
von: Wang, Chao, et al.
Veröffentlicht: (2024)
von: Wang, Chao, et al.
Veröffentlicht: (2024)
NovelQA: Benchmarking Question Answering on Documents Exceeding 200K Tokens
von: Wang, Cunxiang, et al.
Veröffentlicht: (2024)
von: Wang, Cunxiang, et al.
Veröffentlicht: (2024)
Memory-QA: Answering Recall Questions Based on Multimodal Memories
von: Jiang, Hongda, et al.
Veröffentlicht: (2025)
von: Jiang, Hongda, et al.
Veröffentlicht: (2025)
RephQA: Evaluating Readability of Large Language Models in Public Health Question Answering
von: Qiu, Weikang, et al.
Veröffentlicht: (2025)
von: Qiu, Weikang, et al.
Veröffentlicht: (2025)
RetinaQA: A Robust Knowledge Base Question Answering Model for both Answerable and Unanswerable Questions
von: Faldu, Prayushi, et al.
Veröffentlicht: (2024)
von: Faldu, Prayushi, et al.
Veröffentlicht: (2024)
PRIV-QA: Privacy-Preserving Question Answering for Cloud Large Language Models
von: Li, Guangwei, et al.
Veröffentlicht: (2025)
von: Li, Guangwei, et al.
Veröffentlicht: (2025)
M2QA: Multi-domain Multilingual Question Answering
von: Engländer, Leon, et al.
Veröffentlicht: (2024)
von: Engländer, Leon, et al.
Veröffentlicht: (2024)
GRS-QA -- Graph Reasoning-Structured Question Answering Dataset
von: Pahilajani, Anish, et al.
Veröffentlicht: (2024)
von: Pahilajani, Anish, et al.
Veröffentlicht: (2024)
TracSum: A New Benchmark for Aspect-Based Summarization with Sentence-Level Traceability in Medical Domain
von: Chu, Bohao, et al.
Veröffentlicht: (2025)
von: Chu, Bohao, et al.
Veröffentlicht: (2025)
PassiveQA: A Three-Action Framework for Epistemically Calibrated Question Answering via Supervised Finetuning
von: Baidya, Madhav S
Veröffentlicht: (2026)
von: Baidya, Madhav S
Veröffentlicht: (2026)
Do Large Language Models Rank Fairly? An Empirical Study on the Fairness of LLMs as Rankers
von: Wang, Yuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuan, et al.
Veröffentlicht: (2024)
ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2025)
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2025)
FoQA: A Faroese Question-Answering Dataset
von: Simonsen, Annika, et al.
Veröffentlicht: (2025)
von: Simonsen, Annika, et al.
Veröffentlicht: (2025)
SemViQA: A Semantic Question Answering System for Vietnamese Information Fact-Checking
von: Tran, Dien X., et al.
Veröffentlicht: (2025)
von: Tran, Dien X., et al.
Veröffentlicht: (2025)
MMToM-QA: Multimodal Theory of Mind Question Answering
von: Jin, Chuanyang, et al.
Veröffentlicht: (2024)
von: Jin, Chuanyang, et al.
Veröffentlicht: (2024)
ExpliCIT-QA: Explainable Code-Based Image Table Question Answering
von: Lagos, Maximiliano Hormazábal, et al.
Veröffentlicht: (2025)
von: Lagos, Maximiliano Hormazábal, et al.
Veröffentlicht: (2025)
RespondeoQA: a Benchmark for Bilingual Latin-English Question Answering
von: Hudspeth, Marisa, et al.
Veröffentlicht: (2026)
von: Hudspeth, Marisa, et al.
Veröffentlicht: (2026)
DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards
von: Kartha, Aaryaman, et al.
Veröffentlicht: (2025)
von: Kartha, Aaryaman, et al.
Veröffentlicht: (2025)
HistoryBankQA: Multilingual Temporal Question Answering on Historical Events
von: Mandal, Biswadip, et al.
Veröffentlicht: (2025)
von: Mandal, Biswadip, et al.
Veröffentlicht: (2025)
BEnQA: A Question Answering and Reasoning Benchmark for Bengali and English
von: Shafayat, Sheikh, et al.
Veröffentlicht: (2024)
von: Shafayat, Sheikh, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
EnsemHalDet: Robust VLM Hallucination Detection via Ensemble of Internal State Detectors
von: Miyazato, Ryuhei, et al.
Veröffentlicht: (2026) -
Does Reasoning Introduce Bias? A Study of Social Bias Evaluation and Mitigation in LLM Reasoning
von: Wu, Xuyang, et al.
Veröffentlicht: (2025) -
Passage-specific Prompt Tuning for Passage Reranking in Question Answering with Large Language Models
von: Wu, Xuyang, et al.
Veröffentlicht: (2024) -
Table Transformers for Imputing Textual Attributes
von: Wei, Ting-Ruen, et al.
Veröffentlicht: (2024) -
A Survey on Feedback-based Multi-step Reasoning for Large Language Models on Mathematics
von: Wei, Ting-Ruen, et al.
Veröffentlicht: (2025)