DMRetriever: A Family of Models for Improved Text Retrieval in Disaster Management

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yin, Kai, Dong, Xiangjue, Liu, Chengkai, Lin, Allen, Shi, Lingfeng, Mostafavi, Ali, Caverlee, James
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914184439005184
author Yin, Kai
Dong, Xiangjue
Liu, Chengkai
Lin, Allen
Shi, Lingfeng
Mostafavi, Ali
Caverlee, James
author_facet Yin, Kai
Dong, Xiangjue
Liu, Chengkai
Lin, Allen
Shi, Lingfeng
Mostafavi, Ali
Caverlee, James
contents Effective and efficient access to relevant information is essential for disaster management. However, no retrieval model is specialized for disaster management, and existing general-domain models fail to handle the varied search intents inherent to disaster management scenarios, resulting in inconsistent and unreliable performance. To this end, we introduce DMRetriever, the first series of dense retrieval models (33M to 7.6B) tailored for this domain. It is trained through a novel three-stage framework of bidirectional attention adaptation, unsupervised contrastive pre-training, and difficulty-aware progressive instruction fine-tuning, using high-quality data generated through an advanced data refinement pipeline. Comprehensive experiments demonstrate that DMRetriever achieves state-of-the-art (SOTA) performance across all six search intents at every model scale. Moreover, DMRetriever is highly parameter-efficient, with 596M model outperforming baselines over 13.3 X larger and 33M model exceeding baselines with only 7.6% of their parameters. All codes, data, and checkpoints are available at https://github.com/KaiYin97/DMRETRIEVER
format Preprint
id arxiv_https___arxiv_org_abs_2510_15087
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DMRetriever: A Family of Models for Improved Text Retrieval in Disaster Management
Yin, Kai
Dong, Xiangjue
Liu, Chengkai
Lin, Allen
Shi, Lingfeng
Mostafavi, Ali
Caverlee, James
Information Retrieval
Artificial Intelligence
Effective and efficient access to relevant information is essential for disaster management. However, no retrieval model is specialized for disaster management, and existing general-domain models fail to handle the varied search intents inherent to disaster management scenarios, resulting in inconsistent and unreliable performance. To this end, we introduce DMRetriever, the first series of dense retrieval models (33M to 7.6B) tailored for this domain. It is trained through a novel three-stage framework of bidirectional attention adaptation, unsupervised contrastive pre-training, and difficulty-aware progressive instruction fine-tuning, using high-quality data generated through an advanced data refinement pipeline. Comprehensive experiments demonstrate that DMRetriever achieves state-of-the-art (SOTA) performance across all six search intents at every model scale. Moreover, DMRetriever is highly parameter-efficient, with 596M model outperforming baselines over 13.3 X larger and 33M model exceeding baselines with only 7.6% of their parameters. All codes, data, and checkpoints are available at https://github.com/KaiYin97/DMRETRIEVER
title DMRetriever: A Family of Models for Improved Text Retrieval in Disaster Management
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2510.15087