The Art of Scaling Reinforcement Learning Compute for LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Khatri, Devvrit, Madaan, Lovish, Tiwari, Rishabh, Bansal, Rachit, Duvvuri, Sai Surya, Zaheer, Manzil, Dhillon, Inderjit S., Brandfonbrener, David, Agarwal, Rishabh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Interleaved Head Attention
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2026)
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2026)
Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
di: Bansal, Rachit, et al.
Pubblicazione: (2025)
di: Bansal, Rachit, et al.
Pubblicazione: (2025)
Learning, Fast and Slow: Towards LLMs That Adapt Continually
di: Tiwari, Rishabh, et al.
Pubblicazione: (2026)
di: Tiwari, Rishabh, et al.
Pubblicazione: (2026)
LASER: Attention with Exponential Transformation
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2024)
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2024)
LUCID: Attention with Preconditioned Representations
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2026)
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2026)
Rethinking Thinking Tokens: LLMs as Improvement Operators
di: Madaan, Lovish, et al.
Pubblicazione: (2025)
di: Madaan, Lovish, et al.
Pubblicazione: (2025)
Fast and Simplex: 2-Simplicial Attention in Triton
di: Roy, Aurko, et al.
Pubblicazione: (2025)
di: Roy, Aurko, et al.
Pubblicazione: (2025)
Dual-Encoders for Extreme Multi-Label Classification
di: Gupta, Nilesh, et al.
Pubblicazione: (2023)
di: Gupta, Nilesh, et al.
Pubblicazione: (2023)
LoRA Done RITE: Robust Invariant Transformation Equilibration for LoRA Optimization
di: Yen, Jui-Nan, et al.
Pubblicazione: (2024)
di: Yen, Jui-Nan, et al.
Pubblicazione: (2024)
Deep Reinforcement Learning for Sequential Combinatorial Auctions
di: Ravindranath, Sai Srivatsa, et al.
Pubblicazione: (2024)
di: Ravindranath, Sai Srivatsa, et al.
Pubblicazione: (2024)
Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
di: Bansal, Hritik, et al.
Pubblicazione: (2024)
di: Bansal, Hritik, et al.
Pubblicazione: (2024)
Compressing Many-Shots in In-Context Learning
di: Khatri, Devvrit, et al.
Pubblicazione: (2025)
di: Khatri, Devvrit, et al.
Pubblicazione: (2025)
Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data
di: Tang, Yunhao, et al.
Pubblicazione: (2025)
di: Tang, Yunhao, et al.
Pubblicazione: (2025)
Abstract Art Interpretation Using ControlNet
di: Srivastava, Rishabh, et al.
Pubblicazione: (2024)
di: Srivastava, Rishabh, et al.
Pubblicazione: (2024)
Differentially Private Model Merging
di: Yin, Qichuan, et al.
Pubblicazione: (2026)
di: Yin, Qichuan, et al.
Pubblicazione: (2026)
Adaptive Few-Shot Learning (AFSL): Tackling Data Scarcity with Stability, Robustness, and Versatility
di: Agrawal, Rishabh
Pubblicazione: (2025)
di: Agrawal, Rishabh
Pubblicazione: (2025)
A Statistical Framework for Data-dependent Retrieval-Augmented Models
di: Basu, Soumya, et al.
Pubblicazione: (2024)
di: Basu, Soumya, et al.
Pubblicazione: (2024)
ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization
di: Patel, Nirmal, et al.
Pubblicazione: (2026)
di: Patel, Nirmal, et al.
Pubblicazione: (2026)
Geometric Median (GM) Matching for Robust Data Pruning
di: Acharya, Anish, et al.
Pubblicazione: (2024)
di: Acharya, Anish, et al.
Pubblicazione: (2024)
Federation over Text: Insight Sharing for Multi-Agent Reasoning
di: Yao, Dixi, et al.
Pubblicazione: (2026)
di: Yao, Dixi, et al.
Pubblicazione: (2026)
Runtime Evaluation of Procedural Content Generation in an Endless Runner Game Using Autonomous Agents
di: Kar, Rishabh
Pubblicazione: (2026)
di: Kar, Rishabh
Pubblicazione: (2026)
Complement Submodular Information Measures for Balanced and Robust Data Selection
di: Iyer, Rishabh
Pubblicazione: (2026)
di: Iyer, Rishabh
Pubblicazione: (2026)
SiT: Symmetry-Invariant Transformers for Generalisation in Reinforcement Learning
di: Weissenbacher, Matthias, et al.
Pubblicazione: (2024)
di: Weissenbacher, Matthias, et al.
Pubblicazione: (2024)
Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play
di: Ye, Ziyu, et al.
Pubblicazione: (2024)
di: Ye, Ziyu, et al.
Pubblicazione: (2024)
Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers
di: Sareen, Kusha, et al.
Pubblicazione: (2025)
di: Sareen, Kusha, et al.
Pubblicazione: (2025)
Using text embedding models as text classifiers with medical data
di: Goel, Rishabh
Pubblicazione: (2024)
di: Goel, Rishabh
Pubblicazione: (2024)
Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models
di: Huang, Yukun, et al.
Pubblicazione: (2025)
di: Huang, Yukun, et al.
Pubblicazione: (2025)
Using Early Readouts to Mediate Featural Bias in Distillation
di: Tiwari, Rishabh, et al.
Pubblicazione: (2023)
di: Tiwari, Rishabh, et al.
Pubblicazione: (2023)
An Evaluation of Context Length Extrapolation in Long Code via Positional Embeddings and Efficient Attention
di: Ghosh, Madhusudan, et al.
Pubblicazione: (2026)
di: Ghosh, Madhusudan, et al.
Pubblicazione: (2026)
Enhancing Cache-Augmented Generation (CAG) with Adaptive Contextual Compression for Scalable Knowledge Integration
di: Agrawal, Rishabh, et al.
Pubblicazione: (2025)
di: Agrawal, Rishabh, et al.
Pubblicazione: (2025)
SecurePose: Automated Face Blurring and Human Movement Kinematics Extraction from Videos Recorded in Clinical Settings
di: Bajpai, Rishabh, et al.
Pubblicazione: (2024)
di: Bajpai, Rishabh, et al.
Pubblicazione: (2024)
When Dynamics Shift, Robust Task Inference Wins: Offline Imitation Learning with Behavior Foundation Models Revisited
di: Agrawal, Rishabh, et al.
Pubblicazione: (2026)
di: Agrawal, Rishabh, et al.
Pubblicazione: (2026)
Towards Quantifying the Preconditioning Effect of Adam
di: Das, Rudrajit, et al.
Pubblicazione: (2024)
di: Das, Rudrajit, et al.
Pubblicazione: (2024)
Positive Unlabeled Contrastive Learning
di: Acharya, Anish, et al.
Pubblicazione: (2022)
di: Acharya, Anish, et al.
Pubblicazione: (2022)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
Domain-driven Metrics for Reinforcement Learning: A Case Study on Epidemic Control using Agent-based Simulation
di: Gaur, Rishabh, et al.
Pubblicazione: (2025)
di: Gaur, Rishabh, et al.
Pubblicazione: (2025)
Towards Compute-Optimal Many-Shot In-Context Learning
di: Golchin, Shahriar, et al.
Pubblicazione: (2025)
di: Golchin, Shahriar, et al.
Pubblicazione: (2025)
LLM-Guided Monte Carlo Tree Search over Knowledge Graphs: Composing Mechanistic Explanations for Drug-Disease Pairs
di: Jakhar, Rishabh, et al.
Pubblicazione: (2026)
di: Jakhar, Rishabh, et al.
Pubblicazione: (2026)
Scaling Test-Time Compute for Agentic Coding
di: Kim, Joongwon, et al.
Pubblicazione: (2026)
di: Kim, Joongwon, et al.
Pubblicazione: (2026)
Balance Equation-based Distributionally Robust Offline Imitation Learning
di: Agrawal, Rishabh, et al.
Pubblicazione: (2025)
di: Agrawal, Rishabh, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Interleaved Head Attention
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2026) -
Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
di: Bansal, Rachit, et al.
Pubblicazione: (2025) -
Learning, Fast and Slow: Towards LLMs That Adapt Continually
di: Tiwari, Rishabh, et al.
Pubblicazione: (2026) -
LASER: Attention with Exponential Transformation
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2024) -
LUCID: Attention with Preconditioned Representations
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2026)