s3: You Don't Need That Much Data to Train a Search Agent via RL
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Pengcheng, Xu, Xueqiang, Lin, Jiacheng, Xiao, Jinfeng, Wang, Zifeng, Sun, Jimeng, Han, Jiawei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GenRES: Rethinking Evaluation for Generative Relation Extraction in the Era of Large Language Models
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2024)
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2024)
Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2026)
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2026)
TriSum: Learning Summarization Ability from Large Language Models with Structured Rationale
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2024)
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2024)
Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning
von: Lin, Jiacheng, et al.
Veröffentlicht: (2025)
von: Lin, Jiacheng, et al.
Veröffentlicht: (2025)
BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research
von: Wang, Zifeng, et al.
Veröffentlicht: (2025)
von: Wang, Zifeng, et al.
Veröffentlicht: (2025)
Panacea: A foundation model for clinical trial search, summarization, design, and recruitment
von: Lin, Jiacheng, et al.
Veröffentlicht: (2024)
von: Lin, Jiacheng, et al.
Veröffentlicht: (2024)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
von: Tyukin, Georgy, et al.
Veröffentlicht: (2024)
von: Tyukin, Georgy, et al.
Veröffentlicht: (2024)
You Don't Need Prompt Engineering Anymore: The Prompting Inversion
von: Khan, Imran
Veröffentlicht: (2025)
von: Khan, Imran
Veröffentlicht: (2025)
DeepRetrieval: Hacking Real Search Engines and Retrievers with Large Language Models via Reinforcement Learning
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2025)
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2025)
You Don't Need Pre-built Graphs for RAG: Retrieval Augmented Generation with Adaptive Reasoning Structures
von: Chen, Shengyuan, et al.
Veröffentlicht: (2025)
von: Chen, Shengyuan, et al.
Veröffentlicht: (2025)
KG-FIT: Knowledge Graph Fine-Tuning Upon Open-World Knowledge
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2024)
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2024)
Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2024)
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2024)
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
von: Hernandez, Adriano
Veröffentlicht: (2024)
von: Hernandez, Adriano
Veröffentlicht: (2024)
PILOT: Legal Case Outcome Prediction with Case Law
von: Cao, Lang, et al.
Veröffentlicht: (2024)
von: Cao, Lang, et al.
Veröffentlicht: (2024)
Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks
von: Chan, Brian J, et al.
Veröffentlicht: (2024)
von: Chan, Brian J, et al.
Veröffentlicht: (2024)
Zero-Shot Open-Schema Entity Structure Discovery
von: Xu, Xueqiang, et al.
Veröffentlicht: (2025)
von: Xu, Xueqiang, et al.
Veröffentlicht: (2025)
TELEClass: Taxonomy Enrichment and LLM-Enhanced Hierarchical Text Classification with Minimal Supervision
von: Zhang, Yunyi, et al.
Veröffentlicht: (2024)
von: Zhang, Yunyi, et al.
Veröffentlicht: (2024)
MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models
von: Wen, Yilin, et al.
Veröffentlicht: (2023)
von: Wen, Yilin, et al.
Veröffentlicht: (2023)
Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
von: Balepur, Nishant, et al.
Veröffentlicht: (2026)
von: Balepur, Nishant, et al.
Veröffentlicht: (2026)
Your Students Don't Use LLMs Like You Wish They Did
von: Kobler, Sebastian, et al.
Veröffentlicht: (2026)
von: Kobler, Sebastian, et al.
Veröffentlicht: (2026)
Don't Trust Generative Agents to Mimic Communication on Social Networks Unless You Benchmarked their Empirical Realism
von: Münker, Simon, et al.
Veröffentlicht: (2025)
von: Münker, Simon, et al.
Veröffentlicht: (2025)
RAS: Retrieval-And-Structuring for Knowledge-Intensive LLM Generation
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2025)
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2025)
Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing
von: Yuan, Wenhao, et al.
Veröffentlicht: (2026)
von: Yuan, Wenhao, et al.
Veröffentlicht: (2026)
Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval
von: Sinha, Aarush
Veröffentlicht: (2025)
von: Sinha, Aarush
Veröffentlicht: (2025)
Can Large Language Models Replace Data Scientists in Biomedical Research?
von: Wang, Zifeng, et al.
Veröffentlicht: (2024)
von: Wang, Zifeng, et al.
Veröffentlicht: (2024)
Synthetic Data RL: Task Definition Is All You Need
von: Guo, Yiduo, et al.
Veröffentlicht: (2025)
von: Guo, Yiduo, et al.
Veröffentlicht: (2025)
Convomem Benchmark: Why Your First 150 Conversations Don't Need RAG
von: Pakhomov, Egor, et al.
Veröffentlicht: (2025)
von: Pakhomov, Egor, et al.
Veröffentlicht: (2025)
Knowing You Don't Know: Learning When to Continue Search in Multi-round RAG through Self-Practicing
von: Yang, Diji, et al.
Veröffentlicht: (2025)
von: Yang, Diji, et al.
Veröffentlicht: (2025)
Why Don't You Know? Evaluating the Impact of Uncertainty Sources on Uncertainty Quantification in LLMs
von: Goloburda, Maiya, et al.
Veröffentlicht: (2026)
von: Goloburda, Maiya, et al.
Veröffentlicht: (2026)
Vision Transformers Don't Need Trained Registers
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency
von: Wang, Chenlong, et al.
Veröffentlicht: (2025)
von: Wang, Chenlong, et al.
Veröffentlicht: (2025)
Don't Pay Attention
von: Hammoud, Mohammad, et al.
Veröffentlicht: (2025)
von: Hammoud, Mohammad, et al.
Veröffentlicht: (2025)
Thought-Retriever: Don't Just Retrieve Raw Data, Retrieve Thoughts for Memory-Augmented Agentic Systems
von: Feng, Tao, et al.
Veröffentlicht: (2026)
von: Feng, Tao, et al.
Veröffentlicht: (2026)
Don't Touch My Diacritics
von: Gorman, Kyle, et al.
Veröffentlicht: (2024)
von: Gorman, Kyle, et al.
Veröffentlicht: (2024)
Don't exhaust, don't waste
von: Bianchini, Riccardo, et al.
Veröffentlicht: (2025)
von: Bianchini, Riccardo, et al.
Veröffentlicht: (2025)
Sample, Don't Search: Rethinking Test-Time Alignment for Language Models
von: Faria, Gonçalo, et al.
Veröffentlicht: (2025)
von: Faria, Gonçalo, et al.
Veröffentlicht: (2025)
EpiCaR: Knowing What You Don't Know Matters for Better Reasoning in LLMs
von: Yeom, Jewon, et al.
Veröffentlicht: (2026)
von: Yeom, Jewon, et al.
Veröffentlicht: (2026)
Don't Ignore Dual Logic Ability of LLMs while Privatizing: A Data-Intensive Analysis in Medical Domain
von: Du, Yanrui, et al.
Veröffentlicht: (2023)
von: Du, Yanrui, et al.
Veröffentlicht: (2023)
Frictional Agent Alignment Framework: Slow Down and Don't Break Things
von: Nath, Abhijnan, et al.
Veröffentlicht: (2025)
von: Nath, Abhijnan, et al.
Veröffentlicht: (2025)
Don't Throw Away Data: Better Sequence Knowledge Distillation
von: Wang, Jun, et al.
Veröffentlicht: (2024)
von: Wang, Jun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
GenRES: Rethinking Evaluation for Generative Relation Extraction in the Era of Large Language Models
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2024) -
Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2026) -
TriSum: Learning Summarization Ability from Large Language Models with Structured Rationale
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2024) -
Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning
von: Lin, Jiacheng, et al.
Veröffentlicht: (2025) -
BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research
von: Wang, Zifeng, et al.
Veröffentlicht: (2025)