Test-Time Scaling Makes Overtraining Compute-Optimal
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Roberts, Nicholas, Cho, Sungjun, Gao, Zhiqi, Huang, Tzu-Heng, Wu, Albert, Orlanski, Gabriel, Trost, Avi, Buchanan, Kelly, Albarghouthi, Aws, Sala, Frederic |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pareto Optimal Code Generation
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2025)
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2025)
COSMOS: Predictable and Cost-Effective Adaptation of LLMs
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2026)
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2026)
Linear-Time T-Gate Optimization via Random Abstraction
von: Albarghouthi, Aws
Veröffentlicht: (2026)
von: Albarghouthi, Aws
Veröffentlicht: (2026)
SkillOrchestra: Learning to Route Agents via Skill Transfer
von: Wang, Jiayu, et al.
Veröffentlicht: (2026)
von: Wang, Jiayu, et al.
Veröffentlicht: (2026)
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
von: Ge, Albert, et al.
Veröffentlicht: (2025)
von: Ge, Albert, et al.
Veröffentlicht: (2025)
Analyzing Decoders for Quantum Error Correction
von: Molavi, Abtin, et al.
Veröffentlicht: (2026)
von: Molavi, Abtin, et al.
Veröffentlicht: (2026)
Beyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement Learning
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
Weight Updates as Activation Shifts: A Principled Framework for Steering
von: Adila, Dyah, et al.
Veröffentlicht: (2026)
von: Adila, Dyah, et al.
Veröffentlicht: (2026)
A One-Layer Decoder-Only Transformer is a Two-Layer RNN: With an Application to Certified Robustness
von: Zhang, Yuhao, et al.
Veröffentlicht: (2024)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2024)
PECAN: A Deterministic Certified Defense Against Backdoor Attacks
von: Zhang, Yuhao, et al.
Veröffentlicht: (2023)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2023)
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2025)
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2025)
Dependency-Aware Compilation for Surface Code Quantum Architectures
von: Molavi, Abtin, et al.
Veröffentlicht: (2023)
von: Molavi, Abtin, et al.
Veröffentlicht: (2023)
Managing Classical Processing Requirements for Quantum Error Correction
von: Maurya, Satvik, et al.
Veröffentlicht: (2024)
von: Maurya, Satvik, et al.
Veröffentlicht: (2024)
Optimizing Quantum Circuits, Fast and Slow
von: Xu, Amanda, et al.
Veröffentlicht: (2024)
von: Xu, Amanda, et al.
Veröffentlicht: (2024)
ScriptoriumWS: A Code Generation Assistant for Weak Supervision
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2025)
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2025)
Overtrained, Not Misaligned
von: Schreiber, Joel, et al.
Veröffentlicht: (2026)
von: Schreiber, Joel, et al.
Veröffentlicht: (2026)
Reusing Overtrained Language Models Saturates Scaling
von: Liew, Seng Pei, et al.
Veröffentlicht: (2025)
von: Liew, Seng Pei, et al.
Veröffentlicht: (2025)
LLM-Integrated Bayesian State Space Models for Multimodal Time-Series Forecasting
von: Cho, Sungjun, et al.
Veröffentlicht: (2025)
von: Cho, Sungjun, et al.
Veröffentlicht: (2025)
Crowdsourcing Task Traces for Service Robotics
von: Porfirio, David, et al.
Veröffentlicht: (2024)
von: Porfirio, David, et al.
Veröffentlicht: (2024)
Generating Compilers for Qubit Mapping and Routing
von: Molavi, Abtin, et al.
Veröffentlicht: (2025)
von: Molavi, Abtin, et al.
Veröffentlicht: (2025)
Verified Training for Counterfactual Explanation Robustness under Data Shift
von: Meyer, Anna P., et al.
Veröffentlicht: (2024)
von: Meyer, Anna P., et al.
Veröffentlicht: (2024)
Reference-Specific Unlearning Metrics Can Hide the Truth: A Reality Check
von: Cho, Sungjun, et al.
Veröffentlicht: (2025)
von: Cho, Sungjun, et al.
Veröffentlicht: (2025)
LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
Perceptions of the Fairness Impacts of Multiplicity in Machine Learning
von: Meyer, Anna P., et al.
Veröffentlicht: (2024)
von: Meyer, Anna P., et al.
Veröffentlicht: (2024)
The ALCHEmist: Automated Labeling 500x CHEaper Than LLM Data Annotators
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2024)
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2024)
Two-View Accumulation as the Primary Training Lever for Hybrid-Capture Gaussian Splatting: A Variance-Decomposition View of When Gradient Surgery Helps
von: Cho, Sungjun
Veröffentlicht: (2026)
von: Cho, Sungjun
Veröffentlicht: (2026)
Diabetic Retinopathy Grading with CLIP-based Ranking-Aware Adaptation:A Comparative Study on Fundus Image
von: Cho, Sungjun
Veröffentlicht: (2026)
von: Cho, Sungjun
Veröffentlicht: (2026)
U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning
von: Lee, Christine P, et al.
Veröffentlicht: (2026)
von: Lee, Christine P, et al.
Veröffentlicht: (2026)
MoRe Fine-Tuning with 10x Fewer Parameters
von: Tan, Wenxuan, et al.
Veröffentlicht: (2024)
von: Tan, Wenxuan, et al.
Veröffentlicht: (2024)
TARDIS: Mitigating Temporal Misalignment via Representation Steering
von: Shin, Changho, et al.
Veröffentlicht: (2025)
von: Shin, Changho, et al.
Veröffentlicht: (2025)
Models Can Model, But Can't Bind: Structured Grounding in Text-to-Optimization
von: Gao, Zhiqi, et al.
Veröffentlicht: (2026)
von: Gao, Zhiqi, et al.
Veröffentlicht: (2026)
Compute Optimal Scaling of Skills: Knowledge vs Reasoning
von: Roberts, Nicholas, et al.
Veröffentlicht: (2025)
von: Roberts, Nicholas, et al.
Veröffentlicht: (2025)
Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2025)
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2025)
RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2026)
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2026)
Overtrained Language Models Are Harder to Fine-Tune
von: Springer, Jacob Mitchell, et al.
Veröffentlicht: (2025)
von: Springer, Jacob Mitchell, et al.
Veröffentlicht: (2025)
(How) Learning Rates Regulate Catastrophic Overtraining
von: Rofin, Mark, et al.
Veröffentlicht: (2026)
von: Rofin, Mark, et al.
Veröffentlicht: (2026)
Test-time Scaling Techniques in Theoretical Physics -- A Comparison of Methods on the TPBench Dataset
von: Gao, Zhiqi, et al.
Veröffentlicht: (2025)
von: Gao, Zhiqi, et al.
Veröffentlicht: (2025)
Multimodal Data Curation via Object Detection and Filter Ensembles
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2024)
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2024)
Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling
von: Chen, Hao Mark, et al.
Veröffentlicht: (2025)
von: Chen, Hao Mark, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Pareto Optimal Code Generation
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2025) -
COSMOS: Predictable and Cost-Effective Adaptation of LLMs
von: Wang, Jiayu, et al.
Veröffentlicht: (2025) -
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2026) -
Linear-Time T-Gate Optimization via Random Abstraction
von: Albarghouthi, Aws
Veröffentlicht: (2026) -
SkillOrchestra: Learning to Route Agents via Skill Transfer
von: Wang, Jiayu, et al.
Veröffentlicht: (2026)