A Question Answering Dataset for Temporal-Sensitive Retrieval-Augmented Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Ziyang, Min, Erxue, Zhao, Xiang, Li, Yunxin, Jia, Xin, Liao, Jinzhi, Li, Jichao, Wang, Shuaiqiang, Hu, Baotian, Yin, Dawei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908492698222592
author Chen, Ziyang
Min, Erxue
Zhao, Xiang
Li, Yunxin
Jia, Xin
Liao, Jinzhi
Li, Jichao
Wang, Shuaiqiang
Hu, Baotian
Yin, Dawei
author_facet Chen, Ziyang
Min, Erxue
Zhao, Xiang
Li, Yunxin
Jia, Xin
Liao, Jinzhi
Li, Jichao
Wang, Shuaiqiang
Hu, Baotian
Yin, Dawei
contents We introduce ChronoQA, a large-scale benchmark dataset for Chinese question answering, specifically designed to evaluate temporal reasoning in Retrieval-Augmented Generation (RAG) systems. ChronoQA is constructed from over 300,000 news articles published between 2019 and 2024, and contains 5,176 high-quality questions covering absolute, aggregate, and relative temporal types with both explicit and implicit time expressions. The dataset supports both single- and multi-document scenarios, reflecting the real-world requirements for temporal alignment and logical consistency. ChronoQA features comprehensive structural annotations and has undergone multi-stage validation, including rule-based, LLM-based, and human evaluation, to ensure data quality. By providing a dynamic, reliable, and scalable resource, ChronoQA enables structured evaluation across a wide range of temporal tasks, and serves as a robust benchmark for advancing time-sensitive retrieval-augmented question answering systems.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12282
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Question Answering Dataset for Temporal-Sensitive Retrieval-Augmented Generation
Chen, Ziyang
Min, Erxue
Zhao, Xiang
Li, Yunxin
Jia, Xin
Liao, Jinzhi
Li, Jichao
Wang, Shuaiqiang
Hu, Baotian
Yin, Dawei
Computation and Language
Information Retrieval
68T50, 68P20
I.2.7; H.3.3
We introduce ChronoQA, a large-scale benchmark dataset for Chinese question answering, specifically designed to evaluate temporal reasoning in Retrieval-Augmented Generation (RAG) systems. ChronoQA is constructed from over 300,000 news articles published between 2019 and 2024, and contains 5,176 high-quality questions covering absolute, aggregate, and relative temporal types with both explicit and implicit time expressions. The dataset supports both single- and multi-document scenarios, reflecting the real-world requirements for temporal alignment and logical consistency. ChronoQA features comprehensive structural annotations and has undergone multi-stage validation, including rule-based, LLM-based, and human evaluation, to ensure data quality. By providing a dynamic, reliable, and scalable resource, ChronoQA enables structured evaluation across a wide range of temporal tasks, and serves as a robust benchmark for advancing time-sensitive retrieval-augmented question answering systems.
title A Question Answering Dataset for Temporal-Sensitive Retrieval-Augmented Generation
topic Computation and Language
Information Retrieval
68T50, 68P20
I.2.7; H.3.3
url https://arxiv.org/abs/2508.12282