Enregistré dans:
| Auteurs principaux: | Snell, Charlie, Lee, Jaehoon, Xu, Kelvin, Kumar, Aviral |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2408.03314 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
par: Setlur, Amrith, et autres
Publié: (2025)
par: Setlur, Amrith, et autres
Publié: (2025)
Scaling Test-Time Compute Without Verification or RL is Suboptimal
par: Setlur, Amrith, et autres
Publié: (2025)
par: Setlur, Amrith, et autres
Publié: (2025)
Test-Time Scaling Makes Overtraining Compute-Optimal
par: Roberts, Nicholas, et autres
Publié: (2026)
par: Roberts, Nicholas, et autres
Publié: (2026)
Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling
par: Monjur, Ocean, et autres
Publié: (2026)
par: Monjur, Ocean, et autres
Publié: (2026)
RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold
par: Setlur, Amrith, et autres
Publié: (2024)
par: Setlur, Amrith, et autres
Publié: (2024)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
par: Setlur, Amrith, et autres
Publié: (2024)
par: Setlur, Amrith, et autres
Publié: (2024)
Scaling Test-Time Compute for Agentic Coding
par: Kim, Joongwon, et autres
Publié: (2026)
par: Kim, Joongwon, et autres
Publié: (2026)
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
par: Zhao, James Xu, et autres
Publié: (2025)
par: Zhao, James Xu, et autres
Publié: (2025)
Predicting Emergent Capabilities by Finetuning
par: Snell, Charlie, et autres
Publié: (2024)
par: Snell, Charlie, et autres
Publié: (2024)
Value-Based Deep RL Scales Predictably
par: Rybkin, Oleh, et autres
Publié: (2025)
par: Rybkin, Oleh, et autres
Publié: (2025)
Test-Time Scaling with Reflective Generative Model
par: Wang, Zixiao, et autres
Publié: (2025)
par: Wang, Zixiao, et autres
Publié: (2025)
Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
par: Qu, Yuxiao, et autres
Publié: (2025)
par: Qu, Yuxiao, et autres
Publié: (2025)
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge
par: Chan, Chi-Min, et autres
Publié: (2025)
par: Chan, Chi-Min, et autres
Publié: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
par: Zhou, Yilun, et autres
Publié: (2025)
par: Zhou, Yilun, et autres
Publié: (2025)
Resolving Discrepancies in Compute-Optimal Scaling of Language Models
par: Porian, Tomer, et autres
Publié: (2024)
par: Porian, Tomer, et autres
Publié: (2024)
Provable Scaling Laws for the Test-Time Compute of Large Language Models
par: Chen, Yanxi, et autres
Publié: (2024)
par: Chen, Yanxi, et autres
Publié: (2024)
Atom of Thoughts for Markov LLM Test-Time Scaling
par: Teng, Fengwei, et autres
Publié: (2025)
par: Teng, Fengwei, et autres
Publié: (2025)
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
par: MiniMax, et autres
Publié: (2025)
par: MiniMax, et autres
Publié: (2025)
MetaScale: Test-Time Scaling with Evolving Meta-Thoughts
par: Liu, Qin, et autres
Publié: (2025)
par: Liu, Qin, et autres
Publié: (2025)
Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models
par: Wang, Jian, et autres
Publié: (2025)
par: Wang, Jian, et autres
Publié: (2025)
Inference Scaling vs Reasoning: An Empirical Analysis of Compute-Optimal LLM Problem-Solving
par: AbdElhameed, Marwan, et autres
Publié: (2024)
par: AbdElhameed, Marwan, et autres
Publié: (2024)
HEART: Emotionally-Driven Test-Time Scaling of Language Models
par: Pinto, Gabriela, et autres
Publié: (2025)
par: Pinto, Gabriela, et autres
Publié: (2025)
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
par: Geiping, Jonas, et autres
Publié: (2025)
par: Geiping, Jonas, et autres
Publié: (2025)
Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight Models
par: Samadi, Mehrzad, et autres
Publié: (2025)
par: Samadi, Mehrzad, et autres
Publié: (2025)
Kinetics: Rethinking Test-Time Scaling Laws
par: Sadhukhan, Ranajoy, et autres
Publié: (2025)
par: Sadhukhan, Ranajoy, et autres
Publié: (2025)
Log-Augmented Generation: Scaling Test-Time Reasoning with Reusable Computation
par: Chen, Peter Baile, et autres
Publié: (2025)
par: Chen, Peter Baile, et autres
Publié: (2025)
Strategic Scaling of Test-Time Compute: A Bandit Learning Approach
par: Zuo, Bowen, et autres
Publié: (2025)
par: Zuo, Bowen, et autres
Publié: (2025)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
par: Zeng, Zhiyuan, et autres
Publié: (2025)
par: Zeng, Zhiyuan, et autres
Publié: (2025)
RFG: Test-Time Scaling for Diffusion Large Language Model Reasoning with Reward-Free Guidance
par: Chen, Tianlang, et autres
Publié: (2025)
par: Chen, Tianlang, et autres
Publié: (2025)
BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
par: Ding, Dujian, et autres
Publié: (2025)
par: Ding, Dujian, et autres
Publié: (2025)
Rethinking the Role of Prompting Strategies in LLM Test-Time Scaling: A Perspective of Probability Theory
par: Liu, Yexiang, et autres
Publié: (2025)
par: Liu, Yexiang, et autres
Publié: (2025)
MarkovScale: Towards Optimal Sequential Scaling at Inference Time
par: Wang, Youkang, et autres
Publié: (2026)
par: Wang, Youkang, et autres
Publié: (2026)
Parallel Test-Time Scaling for Latent Reasoning Models
par: You, Runyang, et autres
Publié: (2025)
par: You, Runyang, et autres
Publié: (2025)
Prompting Test-Time Scaling Is A Strong LLM Reasoning Data Augmentation
par: Bsharat, Sondos Mahmoud, et autres
Publié: (2025)
par: Bsharat, Sondos Mahmoud, et autres
Publié: (2025)
More Vulnerable than You Think: On the Stability of Tool-Integrated LLM Agents
par: Xiong, Weimin, et autres
Publié: (2025)
par: Xiong, Weimin, et autres
Publié: (2025)
Sleep-time Compute: Beyond Inference Scaling at Test-time
par: Lin, Kevin, et autres
Publié: (2025)
par: Lin, Kevin, et autres
Publié: (2025)
Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences
par: Karlekar, Sweta, et autres
Publié: (2026)
par: Karlekar, Sweta, et autres
Publié: (2026)
Are More Tokens Rational? Inference-Time Scaling in Language Models as Adaptive Resource Rationality
par: Hu, Zhimin, et autres
Publié: (2026)
par: Hu, Zhimin, et autres
Publié: (2026)
Compute Optimal Scaling of Skills: Knowledge vs Reasoning
par: Roberts, Nicholas, et autres
Publié: (2025)
par: Roberts, Nicholas, et autres
Publié: (2025)
SOLAR 10.7B: Scaling Large Language Models with Simple yet Effective Depth Up-Scaling
par: Kim, Dahyun, et autres
Publié: (2023)
par: Kim, Dahyun, et autres
Publié: (2023)
Documents similaires
-
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
par: Setlur, Amrith, et autres
Publié: (2025) -
Scaling Test-Time Compute Without Verification or RL is Suboptimal
par: Setlur, Amrith, et autres
Publié: (2025) -
Test-Time Scaling Makes Overtraining Compute-Optimal
par: Roberts, Nicholas, et autres
Publié: (2026) -
Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling
par: Monjur, Ocean, et autres
Publié: (2026) -
RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold
par: Setlur, Amrith, et autres
Publié: (2024)