DSL-R1: From SQL to DSL for Training Retrieval Agents across Structured and Unstructured Data with Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911534864662528 |
|---|---|
| author | Hu, Yunhai Zhou, Junwei Cao, Yumo Long, Yitao Xu, Yiwei Jiang, Qiyi Wang, Weiyao Cao, Xiaoyu Sun, Zhen Zou, Yiran Du, Nan |
| author_facet | Hu, Yunhai Zhou, Junwei Cao, Yumo Long, Yitao Xu, Yiwei Jiang, Qiyi Wang, Weiyao Cao, Xiaoyu Sun, Zhen Zou, Yiran Du, Nan |
| contents | Effective retrieval in complex domains requires bridging the gap between structured metadata and unstructured content. Existing systems typically isolate these capabilities, relying on either symbolic filtering or vector similarity, failing to capture their interplay. In this work, we propose DSL-R1, a unified framework that synergizes logical reasoning with semantic matching via a novel Domain-Specific Language (DSL). By embedding vector primitives within SQL-style operators, our approach leverages the complementary strengths of symbolic precision and semantic coverage. We further introduce a reinforcement learning mechanism where rule-based execution feedback and retrieval quality rewards jointly optimize the DSL generation, balancing structural correctness and semantic alignment. Evaluations on a large-scale industrial email benchmark demonstrate that DSL-R1 achieves a +12.3% improvement in Hit@1/3, consistently outperforming decoupled baselines and establishing a robust paradigm for hybrid retrieval. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_21018 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | DSL-R1: From SQL to DSL for Training Retrieval Agents across Structured and Unstructured Data with Reinforcement Learning Hu, Yunhai Zhou, Junwei Cao, Yumo Long, Yitao Xu, Yiwei Jiang, Qiyi Wang, Weiyao Cao, Xiaoyu Sun, Zhen Zou, Yiran Du, Nan Information Retrieval Artificial Intelligence Databases Machine Learning Effective retrieval in complex domains requires bridging the gap between structured metadata and unstructured content. Existing systems typically isolate these capabilities, relying on either symbolic filtering or vector similarity, failing to capture their interplay. In this work, we propose DSL-R1, a unified framework that synergizes logical reasoning with semantic matching via a novel Domain-Specific Language (DSL). By embedding vector primitives within SQL-style operators, our approach leverages the complementary strengths of symbolic precision and semantic coverage. We further introduce a reinforcement learning mechanism where rule-based execution feedback and retrieval quality rewards jointly optimize the DSL generation, balancing structural correctness and semantic alignment. Evaluations on a large-scale industrial email benchmark demonstrate that DSL-R1 achieves a +12.3% improvement in Hit@1/3, consistently outperforming decoupled baselines and establishing a robust paradigm for hybrid retrieval. |
| title | DSL-R1: From SQL to DSL for Training Retrieval Agents across Structured and Unstructured Data with Reinforcement Learning |
| topic | Information Retrieval Artificial Intelligence Databases Machine Learning |
| url | https://arxiv.org/abs/2603.21018 |