PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost
Fuente:
arXiv
Saved in:
| Main Authors: | Yi, Junkeun, Mosk-Aoyama, Damon, Huang, Baihe, Gala, Ritu, Wang, Charles, Devare, Sugam Dipak, Bhardwaj, Khushi, Gupta, Abhibha, Kuchaiev, Oleksii, Jiao, Jiantao, Zhang, Jian, Srinivasan, Venkat |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Bits to Rounds: Parallel Decoding with Exploration for Diffusion Language Models
by: Fu, Hengyu, et al.
Published: (2025)
by: Fu, Hengyu, et al.
Published: (2025)
GPT vs RETRO: Exploring the Intersection of Retrieval and Parameter-Efficient Fine-Tuning
by: Ficek, Aleksander, et al.
Published: (2024)
by: Ficek, Aleksander, et al.
Published: (2024)
Tied-Lora: Enhancing parameter efficiency of LoRA with weight tying
by: Renduchintala, Adithya, et al.
Published: (2023)
by: Renduchintala, Adithya, et al.
Published: (2023)
Towards Anytime-Valid Statistical Watermarking
by: Huang, Baihe, et al.
Published: (2026)
by: Huang, Baihe, et al.
Published: (2026)
Think Twice: Branch-and-Rethink Reasoning Reward Model
by: Jiao, Yizhu, et al.
Published: (2025)
by: Jiao, Yizhu, et al.
Published: (2025)
Structured Context Engineering for File-Native Agentic Systems: Evaluating Schema Accuracy, Format Effectiveness, and Multi-File Navigation at Scale
by: McMillan, Damon
Published: (2026)
by: McMillan, Damon
Published: (2026)
Catalytic Hydrogenation of CO2 by Direct Air Capture to Valuable C1 Products Using Homogenous Catalysts
by: Ritu Bhardwaj, et al.
Published: (2025)
by: Ritu Bhardwaj, et al.
Published: (2025)
Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics
by: Zhu, Hanlin, et al.
Published: (2024)
by: Zhu, Hanlin, et al.
Published: (2024)
Pre-training LLMs using human-like development data corpus
by: Bhardwaj, Khushi, et al.
Published: (2023)
by: Bhardwaj, Khushi, et al.
Published: (2023)
Development of Cognitive Intelligence in Pre-trained Language Models
by: Shah, Raj Sanjay, et al.
Published: (2024)
by: Shah, Raj Sanjay, et al.
Published: (2024)
Transmission matrix measurement of a single Mie scatterer
by: Sui, Xiaomeng, et al.
Published: (2026)
by: Sui, Xiaomeng, et al.
Published: (2026)
COMPOST MANURE PREPARATION AND PEST MANAGEMENT IN ORGANIC FARMING
by: Sunil, Dahal, et al.
Published: (2025)
by: Sunil, Dahal, et al.
Published: (2025)
Same Ranking, Different Winner: How Scoring Targets Shape LLM Memory Benchmarks
by: Panthi, Sugam, et al.
Published: (2026)
by: Panthi, Sugam, et al.
Published: (2026)
Towards Optimal Statistical Watermarking
by: Huang, Baihe, et al.
Published: (2023)
by: Huang, Baihe, et al.
Published: (2023)
Sample Complexity and Representation Ability of Test-time Scaling Paradigms
by: Huang, Baihe, et al.
Published: (2025)
by: Huang, Baihe, et al.
Published: (2025)
Better Synthetic Data by Retrieving and Transforming Existing Datasets
by: Gandhi, Saumya, et al.
Published: (2024)
by: Gandhi, Saumya, et al.
Published: (2024)
Extracting OPQRST in Electronic Health Records using Large Language Models with Reasoning
by: Luo, Zhimeng, et al.
Published: (2025)
by: Luo, Zhimeng, et al.
Published: (2025)
Contracts: A unified lens on congestion control robustness, fairness, congestion, and generality
by: Agarwal, Anup, et al.
Published: (2025)
by: Agarwal, Anup, et al.
Published: (2025)
GSM-Agent: Understanding Agentic Reasoning Using Controllable Environments
by: Zhu, Hanlin, et al.
Published: (2025)
by: Zhu, Hanlin, et al.
Published: (2025)
Labor Laws and Workers Rights in India
by: Khushi Goyal
Published: (2026)
by: Khushi Goyal
Published: (2026)
INSIDER TRADING AND SEBI REGULATIONS
by: Khushi Rani
Published: (2026)
by: Khushi Rani
Published: (2026)
A FIELD REPORT ON THE CHILD MARRIAGE IN GAGRI VILLAGE SEIKHPURA, BIHAR
by: Khushi Rani
Published: (2026)
by: Khushi Rani
Published: (2026)
ERA of Development in Jammu and Kashmir After the Abrogation of Article 370
by: Khushi Gupta
Published: (2025)
by: Khushi Gupta
Published: (2025)
Leveraging Technology, to Enhance Workplace Communication: Impact, Challenges, and Best Practices
by: Khushi Malviya
Published: (2025)
by: Khushi Malviya
Published: (2025)
Noise-robust latent vector reconstruction in ptychography using deep generative models
by: Seifert, Jacob, et al.
Published: (2023)
by: Seifert, Jacob, et al.
Published: (2023)
Multiparameter Maximum Information States for Coherent Diffraction Measurements
by: Verreussel, Bram, et al.
Published: (2026)
by: Verreussel, Bram, et al.
Published: (2026)
Formal Analysis and Supply Chain Security for Agentic AI Skills
by: Bhardwaj, Varun Pratap
Published: (2026)
by: Bhardwaj, Varun Pratap
Published: (2026)
Quantifying the Accuracy and Cost Impact of Design Decisions in Budget-Constrained Agentic LLM Search
by: McCleary, Kyle, et al.
Published: (2026)
by: McCleary, Kyle, et al.
Published: (2026)
Human Behavioral Benchmarking: Numeric Magnitude Comparison Effects in Large Language Models
by: Shah, Raj Sanjay, et al.
Published: (2023)
by: Shah, Raj Sanjay, et al.
Published: (2023)
HelpSteer2-Preference: Complementing Ratings with Preferences
by: Wang, Zhilin, et al.
Published: (2024)
by: Wang, Zhilin, et al.
Published: (2024)
SoK: Agentic Skills -- Beyond Tool Use in LLM Agents
by: Jiang, Yanna, et al.
Published: (2026)
by: Jiang, Yanna, et al.
Published: (2026)
PiCA: Pivot-Based Credit Assignment for Search Agentic Reinforcement Learning
by: Liu, Dongyi, et al.
Published: (2026)
by: Liu, Dongyi, et al.
Published: (2026)
Toward a Theory of Tokenization in LLMs
by: Rajaraman, Nived, et al.
Published: (2024)
by: Rajaraman, Nived, et al.
Published: (2024)
Efficient Prompt Caching via Embedding Similarity
by: Zhu, Hanlin, et al.
Published: (2024)
by: Zhu, Hanlin, et al.
Published: (2024)
Rate-Cost Tradeoffs in Nonlinear Control
by: Atay, Eray Unsal, et al.
Published: (2026)
by: Atay, Eray Unsal, et al.
Published: (2026)
Skill Reuse as Compression in Agentic RL
by: Xu, Zhikun, et al.
Published: (2026)
by: Xu, Zhikun, et al.
Published: (2026)
On the Performance of an Explainable Language Model on PubMedQA
by: Srinivasan, Venkat, et al.
Published: (2025)
by: Srinivasan, Venkat, et al.
Published: (2025)
Gyan: An Explainable Neuro-Symbolic Language Model
by: Srinivasan, Venkat, et al.
Published: (2026)
by: Srinivasan, Venkat, et al.
Published: (2026)
The Costs of Early-career Disciplinary Pivots: Evidence from Ph.D. Admissions
by: Xiang, Sidney, et al.
Published: (2026)
by: Xiang, Sidney, et al.
Published: (2026)
Adversarial Training of Reward Models
by: Bukharin, Alexander, et al.
Published: (2025)
by: Bukharin, Alexander, et al.
Published: (2025)
Similar Items
-
From Bits to Rounds: Parallel Decoding with Exploration for Diffusion Language Models
by: Fu, Hengyu, et al.
Published: (2025) -
GPT vs RETRO: Exploring the Intersection of Retrieval and Parameter-Efficient Fine-Tuning
by: Ficek, Aleksander, et al.
Published: (2024) -
Tied-Lora: Enhancing parameter efficiency of LoRA with weight tying
by: Renduchintala, Adithya, et al.
Published: (2023) -
Towards Anytime-Valid Statistical Watermarking
by: Huang, Baihe, et al.
Published: (2026) -
Think Twice: Branch-and-Rethink Reasoning Reward Model
by: Jiao, Yizhu, et al.
Published: (2025)