DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shao, Rulin, Asai, Akari, Shen, Shannon Zejiang, Ivison, Hamish, Kishore, Varsha, Zhuo, Jingming, Zhao, Xinran, Park, Molly, Finlayson, Samuel G., Sontag, David, Murray, Tyler, Min, Sewon, Dasigi, Pradeep, Soldaini, Luca, Brahman, Faeze, Yih, Wen-tau, Wu, Tongshuang, Zettlemoyer, Luke, Kim, Yoon, Hajishirzi, Hannaneh, Koh, Pang Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916014832222208
author Shao, Rulin
Asai, Akari
Shen, Shannon Zejiang
Ivison, Hamish
Kishore, Varsha
Zhuo, Jingming
Zhao, Xinran
Park, Molly
Finlayson, Samuel G.
Sontag, David
Murray, Tyler
Min, Sewon
Dasigi, Pradeep
Soldaini, Luca
Brahman, Faeze
Yih, Wen-tau
Wu, Tongshuang
Zettlemoyer, Luke
Kim, Yoon
Hajishirzi, Hannaneh
Koh, Pang Wei
author_facet Shao, Rulin
Asai, Akari
Shen, Shannon Zejiang
Ivison, Hamish
Kishore, Varsha
Zhuo, Jingming
Zhao, Xinran
Park, Molly
Finlayson, Samuel G.
Sontag, David
Murray, Tyler
Min, Sewon
Dasigi, Pradeep
Soldaini, Luca
Brahman, Faeze
Yih, Wen-tau
Wu, Tongshuang
Zettlemoyer, Luke
Kim, Yoon
Hajishirzi, Hannaneh
Koh, Pang Wei
contents Deep research agents perform multi-step research to produce long-form, well-attributed answers. However, most open deep research agents are trained on easily verifiable short-form QA tasks via reinforcement learning with verifiable rewards, which does not extend to realistic long-form tasks. We address this with Reinforcement Learning with Evolving Rubrics (RLER), where rubrics are constructed and maintained to co-evolve with the policy model during training. This allows the rubrics to incorporate newly explored information from search and contrasting model responses, enabling better fact checking and more discriminative on-policy feedback. Using RLER, we develop Deep Research Tulu (DR Tulu-8B), the first fully open model that is directly trained for open-ended, long-form deep research. Across four long-form deep research benchmarks in science, healthcare, and general domains, DR Tulu substantially outperforms existing open deep research agents (by 15.6% over Tongyi DR on average) and matches or exceeds proprietary deep research agents (by 0.7% over OpenAI DR on average), while being significantly smaller and cheaper per query (1000x cheaper than OpenAI DR per query).
format Preprint
id arxiv_https___arxiv_org_abs_2511_19399
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
Shao, Rulin
Asai, Akari
Shen, Shannon Zejiang
Ivison, Hamish
Kishore, Varsha
Zhuo, Jingming
Zhao, Xinran
Park, Molly
Finlayson, Samuel G.
Sontag, David
Murray, Tyler
Min, Sewon
Dasigi, Pradeep
Soldaini, Luca
Brahman, Faeze
Yih, Wen-tau
Wu, Tongshuang
Zettlemoyer, Luke
Kim, Yoon
Hajishirzi, Hannaneh
Koh, Pang Wei
Computation and Language
Artificial Intelligence
Machine Learning
Deep research agents perform multi-step research to produce long-form, well-attributed answers. However, most open deep research agents are trained on easily verifiable short-form QA tasks via reinforcement learning with verifiable rewards, which does not extend to realistic long-form tasks. We address this with Reinforcement Learning with Evolving Rubrics (RLER), where rubrics are constructed and maintained to co-evolve with the policy model during training. This allows the rubrics to incorporate newly explored information from search and contrasting model responses, enabling better fact checking and more discriminative on-policy feedback. Using RLER, we develop Deep Research Tulu (DR Tulu-8B), the first fully open model that is directly trained for open-ended, long-form deep research. Across four long-form deep research benchmarks in science, healthcare, and general domains, DR Tulu substantially outperforms existing open deep research agents (by 15.6% over Tongyi DR on average) and matches or exceeds proprietary deep research agents (by 0.7% over OpenAI DR on average), while being significantly smaller and cheaper per query (1000x cheaper than OpenAI DR per query).
title DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2511.19399