CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914124077727744 |
|---|---|
| author | Wang, Jiaqi Yang, Xiao Sun, Kai Suresh, Parth Sharma, Sanat Czyzewski, Adam Andersen, Derek Appini, Surya Banerjee, Arkav Choudhary, Sajal Ghasemlou, Shervin Guan, Ziqiang Iyer, Akil Khan, Haidar Kong, Lingkun Luo, Roy Ma, Tiffany Qiao, Zhen Tran, David Xu, Wenfang Yeatman, Skyler Zhou, Chen Gujral, Gunveer Xia, Yinglong Moon, Shane Scheffer, Nicolas Shah, Nirav Chang, Eun Liu, Yue Metze, Florian Stark, Tammy Feizollahi, Zhaleh Jessee, Andrea Pujari, Mangesh Aly, Ahmed Damavandi, Babak Wanga, Rakesh Kumar, Anuj Patel, Rohit Yih, Wen-tau Dong, Xin Luna |
| author_facet | Wang, Jiaqi Yang, Xiao Sun, Kai Suresh, Parth Sharma, Sanat Czyzewski, Adam Andersen, Derek Appini, Surya Banerjee, Arkav Choudhary, Sajal Ghasemlou, Shervin Guan, Ziqiang Iyer, Akil Khan, Haidar Kong, Lingkun Luo, Roy Ma, Tiffany Qiao, Zhen Tran, David Xu, Wenfang Yeatman, Skyler Zhou, Chen Gujral, Gunveer Xia, Yinglong Moon, Shane Scheffer, Nicolas Shah, Nirav Chang, Eun Liu, Yue Metze, Florian Stark, Tammy Feizollahi, Zhaleh Jessee, Andrea Pujari, Mangesh Aly, Ahmed Damavandi, Babak Wanga, Rakesh Kumar, Anuj Patel, Rohit Yih, Wen-tau Dong, Xin Luna |
| contents | Wearable devices such as smart glasses are transforming the way people interact with their surroundings, enabling users to seek information regarding entities in their view. Multi-Modal Retrieval-Augmented Generation (MM-RAG) plays a key role in supporting such questions, yet there is still no comprehensive benchmark for this task, especially regarding wearables scenarios. To fill this gap, we present CRAG-MM -- a Comprehensive RAG benchmark for Multi-modal Multi-turn conversations. CRAG-MM contains a diverse set of 6.5K (image, question, answer) triplets and 2K visual-based multi-turn conversations across 13 domains, including 6.2K egocentric images designed to mimic captures from wearable devices. We carefully constructed the questions to reflect real-world scenarios and challenges, including five types of image-quality issues, six question types, varying entity popularity, differing information dynamism, and different conversation turns. We design three tasks: single-source augmentation, multi-source augmentation, and multi-turn conversations -- each paired with an associated retrieval corpus and APIs for both image-KG retrieval and webpage retrieval. Our evaluation shows that straightforward RAG approaches achieve only 32% and 43% truthfulness on CRAG-MM single- and multi-turn QA, respectively, whereas state-of-the-art industry solutions have similar quality (32%/45%), underscoring ample room for improvement. The benchmark has hosted KDD Cup 2025, attracting about 1K participants and 5K submissions, with winning solutions improving baseline performance by 28%, highlighting its early impact on advancing the field. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_26160 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark Wang, Jiaqi Yang, Xiao Sun, Kai Suresh, Parth Sharma, Sanat Czyzewski, Adam Andersen, Derek Appini, Surya Banerjee, Arkav Choudhary, Sajal Ghasemlou, Shervin Guan, Ziqiang Iyer, Akil Khan, Haidar Kong, Lingkun Luo, Roy Ma, Tiffany Qiao, Zhen Tran, David Xu, Wenfang Yeatman, Skyler Zhou, Chen Gujral, Gunveer Xia, Yinglong Moon, Shane Scheffer, Nicolas Shah, Nirav Chang, Eun Liu, Yue Metze, Florian Stark, Tammy Feizollahi, Zhaleh Jessee, Andrea Pujari, Mangesh Aly, Ahmed Damavandi, Babak Wanga, Rakesh Kumar, Anuj Patel, Rohit Yih, Wen-tau Dong, Xin Luna Computer Vision and Pattern Recognition Wearable devices such as smart glasses are transforming the way people interact with their surroundings, enabling users to seek information regarding entities in their view. Multi-Modal Retrieval-Augmented Generation (MM-RAG) plays a key role in supporting such questions, yet there is still no comprehensive benchmark for this task, especially regarding wearables scenarios. To fill this gap, we present CRAG-MM -- a Comprehensive RAG benchmark for Multi-modal Multi-turn conversations. CRAG-MM contains a diverse set of 6.5K (image, question, answer) triplets and 2K visual-based multi-turn conversations across 13 domains, including 6.2K egocentric images designed to mimic captures from wearable devices. We carefully constructed the questions to reflect real-world scenarios and challenges, including five types of image-quality issues, six question types, varying entity popularity, differing information dynamism, and different conversation turns. We design three tasks: single-source augmentation, multi-source augmentation, and multi-turn conversations -- each paired with an associated retrieval corpus and APIs for both image-KG retrieval and webpage retrieval. Our evaluation shows that straightforward RAG approaches achieve only 32% and 43% truthfulness on CRAG-MM single- and multi-turn QA, respectively, whereas state-of-the-art industry solutions have similar quality (32%/45%), underscoring ample room for improvement. The benchmark has hosted KDD Cup 2025, attracting about 1K participants and 5K submissions, with winning solutions improving baseline performance by 28%, highlighting its early impact on advancing the field. |
| title | CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2510.26160 |