ORBIT: Scalable and Verifiable Data Generation for Search Agents on a Tight Budget
Fuente:
arXiv
Saved in:
| Main Authors: | Thakur, Nandan, Chen, Zijian, Ma, Xueguang, Lin, Jimmy |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs
by: Thakur, Nandan, et al.
Published: (2025)
by: Thakur, Nandan, et al.
Published: (2025)
Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks
by: Kuissi, Nathan, et al.
Published: (2026)
by: Kuissi, Nathan, et al.
Published: (2026)
Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval
by: Thakur, Nandan, et al.
Published: (2023)
by: Thakur, Nandan, et al.
Published: (2023)
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
by: Thakur, Nandan, et al.
Published: (2025)
by: Thakur, Nandan, et al.
Published: (2025)
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
by: Thakur, Nandan, et al.
Published: (2025)
by: Thakur, Nandan, et al.
Published: (2025)
Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track
by: Pradeep, Ronak, et al.
Published: (2024)
by: Pradeep, Ronak, et al.
Published: (2024)
An Efficient Rubric-based Generative Verifier for Search-Augmented LLMs
by: Ma, Linyue, et al.
Published: (2025)
by: Ma, Linyue, et al.
Published: (2025)
Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?
by: Hsu, Tz-Huan, et al.
Published: (2026)
by: Hsu, Tz-Huan, et al.
Published: (2026)
LLatrieval: LLM-Verified Retrieval for Verifiable Generation
by: Li, Xiaonan, et al.
Published: (2023)
by: Li, Xiaonan, et al.
Published: (2023)
Argus: Evidence Assembly for Scalable Deep Research Agents
by: Zhang, Zhen, et al.
Published: (2026)
by: Zhang, Zhen, et al.
Published: (2026)
Learning to Ask: Conversational Product Search via Representation Learning
by: Zou, Jie, et al.
Published: (2024)
by: Zou, Jie, et al.
Published: (2024)
A Survey on Retrieval-Augmented Text Generation for Large Language Models
by: Huang, Yizheng, et al.
Published: (2024)
by: Huang, Yizheng, et al.
Published: (2024)
VerifAI: A Verifiable Open-Source Search Engine for Biomedical Question Answering
by: Košprdić, Miloš, et al.
Published: (2026)
by: Košprdić, Miloš, et al.
Published: (2026)
CoSearchAgent: A Lightweight Collaborative Search Agent with Large Language Models
by: Gong, Peiyuan, et al.
Published: (2024)
by: Gong, Peiyuan, et al.
Published: (2024)
Level-Navi Agent: A Framework and benchmark for Chinese Web Search Agents
by: Hu, Chuanrui, et al.
Published: (2024)
by: Hu, Chuanrui, et al.
Published: (2024)
OrgForge: A Multi-Agent Simulation Framework for Verifiable Synthetic Corporate Corpora
by: Flynt, Jeffrey
Published: (2026)
by: Flynt, Jeffrey
Published: (2026)
OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis
by: Li, Zhuofeng, et al.
Published: (2026)
by: Li, Zhuofeng, et al.
Published: (2026)
LACONIC: Dense-Level Effectiveness for Scalable Sparse Retrieval via a Two-Phase Training Curriculum
by: Xu, Zhichao, et al.
Published: (2026)
by: Xu, Zhichao, et al.
Published: (2026)
Rank-R1: Enhancing Reasoning in LLM-based Document Rerankers via Reinforcement Learning
by: Zhuang, Shengyao, et al.
Published: (2025)
by: Zhuang, Shengyao, et al.
Published: (2025)
OneSearch-V2: The Latent Reasoning Enhanced Self-distillation Generative Search Framework
by: Chen, Ben, et al.
Published: (2026)
by: Chen, Ben, et al.
Published: (2026)
BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
by: Chen, Zijian, et al.
Published: (2025)
by: Chen, Zijian, et al.
Published: (2025)
The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models
by: Pradeep, Ronak, et al.
Published: (2025)
by: Pradeep, Ronak, et al.
Published: (2025)
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework
by: Pradeep, Ronak, et al.
Published: (2024)
by: Pradeep, Ronak, et al.
Published: (2024)
DeepAgent: A General Reasoning Agent with Scalable Toolsets
by: Li, Xiaoxi, et al.
Published: (2025)
by: Li, Xiaoxi, et al.
Published: (2025)
ProductAgent: Benchmarking Conversational Product Search Agent with Asking Clarification Questions
by: Ye, Jingheng, et al.
Published: (2024)
by: Ye, Jingheng, et al.
Published: (2024)
A Survey of LLM-based Deep Search Agents: Paradigm, Optimization, Evaluation, and Challenges
by: Xi, Yunjia, et al.
Published: (2025)
by: Xi, Yunjia, et al.
Published: (2025)
Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented Reasoning
by: Liang, Zihan, et al.
Published: (2026)
by: Liang, Zihan, et al.
Published: (2026)
SD-Search: On-Policy Hindsight Self-Distillation for Search-Augmented Reasoning
by: Ma, Yufei, et al.
Published: (2026)
by: Ma, Yufei, et al.
Published: (2026)
TURA: Tool-Augmented Unified Retrieval Agent for AI Search
by: Zhao, Zhejun, et al.
Published: (2025)
by: Zhao, Zhejun, et al.
Published: (2025)
IG-Search: Step-Level Information Gain Rewards for Search-Augmented Reasoning
by: Liang, Zihan, et al.
Published: (2026)
by: Liang, Zihan, et al.
Published: (2026)
Controlling Output Rankings in Generative Engines for LLM-based Search
by: Jin, Haibo, et al.
Published: (2026)
by: Jin, Haibo, et al.
Published: (2026)
Overview of the TREC 2021 deep learning track
by: Craswell, Nick, et al.
Published: (2025)
by: Craswell, Nick, et al.
Published: (2025)
Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses
by: Jiang, Pengcheng, et al.
Published: (2026)
by: Jiang, Pengcheng, et al.
Published: (2026)
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
by: Jin, Bowen, et al.
Published: (2025)
by: Jin, Bowen, et al.
Published: (2025)
Multi-objective Aligned Bidword Generation Model for E-commerce Search Advertising
by: Liu, Zhenhui, et al.
Published: (2025)
by: Liu, Zhenhui, et al.
Published: (2025)
Evaluating Robustness of Generative Search Engine on Adversarial Factual Questions
by: Hu, Xuming, et al.
Published: (2024)
by: Hu, Xuming, et al.
Published: (2024)
Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs
by: Seo, Yongsik, et al.
Published: (2026)
by: Seo, Yongsik, et al.
Published: (2026)
DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense Retrievers
by: Ma, Xueguang, et al.
Published: (2025)
by: Ma, Xueguang, et al.
Published: (2025)
Evaluating Chain-of-Thought Reasoning through Reusability and Verifiability
by: Aggarwal, Shashank, et al.
Published: (2026)
by: Aggarwal, Shashank, et al.
Published: (2026)
OpenNovelty: An LLM-powered Agentic System for Verifiable Scholarly Novelty Assessment
by: Zhang, Ming, et al.
Published: (2026)
by: Zhang, Ming, et al.
Published: (2026)
Similar Items
-
Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs
by: Thakur, Nandan, et al.
Published: (2025) -
Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks
by: Kuissi, Nathan, et al.
Published: (2026) -
Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval
by: Thakur, Nandan, et al.
Published: (2023) -
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
by: Thakur, Nandan, et al.
Published: (2025) -
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
by: Thakur, Nandan, et al.
Published: (2025)