A Question Answering Dataset for Temporal-Sensitive Retrieval-Augmented Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908492698222592 |
|---|---|
| author | Chen, Ziyang Min, Erxue Zhao, Xiang Li, Yunxin Jia, Xin Liao, Jinzhi Li, Jichao Wang, Shuaiqiang Hu, Baotian Yin, Dawei |
| author_facet | Chen, Ziyang Min, Erxue Zhao, Xiang Li, Yunxin Jia, Xin Liao, Jinzhi Li, Jichao Wang, Shuaiqiang Hu, Baotian Yin, Dawei |
| contents | We introduce ChronoQA, a large-scale benchmark dataset for Chinese question answering, specifically designed to evaluate temporal reasoning in Retrieval-Augmented Generation (RAG) systems. ChronoQA is constructed from over 300,000 news articles published between 2019 and 2024, and contains 5,176 high-quality questions covering absolute, aggregate, and relative temporal types with both explicit and implicit time expressions. The dataset supports both single- and multi-document scenarios, reflecting the real-world requirements for temporal alignment and logical consistency. ChronoQA features comprehensive structural annotations and has undergone multi-stage validation, including rule-based, LLM-based, and human evaluation, to ensure data quality. By providing a dynamic, reliable, and scalable resource, ChronoQA enables structured evaluation across a wide range of temporal tasks, and serves as a robust benchmark for advancing time-sensitive retrieval-augmented question answering systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_12282 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Question Answering Dataset for Temporal-Sensitive Retrieval-Augmented Generation Chen, Ziyang Min, Erxue Zhao, Xiang Li, Yunxin Jia, Xin Liao, Jinzhi Li, Jichao Wang, Shuaiqiang Hu, Baotian Yin, Dawei Computation and Language Information Retrieval 68T50, 68P20 I.2.7; H.3.3 We introduce ChronoQA, a large-scale benchmark dataset for Chinese question answering, specifically designed to evaluate temporal reasoning in Retrieval-Augmented Generation (RAG) systems. ChronoQA is constructed from over 300,000 news articles published between 2019 and 2024, and contains 5,176 high-quality questions covering absolute, aggregate, and relative temporal types with both explicit and implicit time expressions. The dataset supports both single- and multi-document scenarios, reflecting the real-world requirements for temporal alignment and logical consistency. ChronoQA features comprehensive structural annotations and has undergone multi-stage validation, including rule-based, LLM-based, and human evaluation, to ensure data quality. By providing a dynamic, reliable, and scalable resource, ChronoQA enables structured evaluation across a wide range of temporal tasks, and serves as a robust benchmark for advancing time-sensitive retrieval-augmented question answering systems. |
| title | A Question Answering Dataset for Temporal-Sensitive Retrieval-Augmented Generation |
| topic | Computation and Language Information Retrieval 68T50, 68P20 I.2.7; H.3.3 |
| url | https://arxiv.org/abs/2508.12282 |