Saved in:
| Main Authors: | Filatov, Oleg, Wang, Jiangtao, Ebert, Jan, Kesselheim, Stefan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2510.03871 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Time Transfer: On Optimal Learning Rate and Batch Size In The Infinite Data Limit
by: Filatov, Oleg, et al.
Published: (2024)
by: Filatov, Oleg, et al.
Published: (2024)
Memory and Bandwidth are All You Need for Fully Sharded Data Parallel
by: Wang, Jiangtao, et al.
Published: (2025)
by: Wang, Jiangtao, et al.
Published: (2025)
OptScale: Probabilistic Optimality for Inference-time Scaling
by: Wang, Youkang, et al.
Published: (2025)
by: Wang, Youkang, et al.
Published: (2025)
Polynomial, trigonometric, and tropical activations
by: Khalfaoui-Hassani, Ismail, et al.
Published: (2025)
by: Khalfaoui-Hassani, Ismail, et al.
Published: (2025)
MarkovScale: Towards Optimal Sequential Scaling at Inference Time
by: Wang, Youkang, et al.
Published: (2026)
by: Wang, Youkang, et al.
Published: (2026)
Compute-Optimal LLMs Provably Generalize Better With Scale
by: Finzi, Marc, et al.
Published: (2025)
by: Finzi, Marc, et al.
Published: (2025)
syftr: Pareto-Optimal Generative AI
by: Conway, Alexander, et al.
Published: (2025)
by: Conway, Alexander, et al.
Published: (2025)
Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time Scaling
by: Wang, Xinglin, et al.
Published: (2025)
by: Wang, Xinglin, et al.
Published: (2025)
Data Pruning in Generative Diffusion Models
by: Briq, Rania, et al.
Published: (2024)
by: Briq, Rania, et al.
Published: (2024)
Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling
by: Chen, Hao Mark, et al.
Published: (2025)
by: Chen, Hao Mark, et al.
Published: (2025)
IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RL
by: Cheng, Zhoujun, et al.
Published: (2026)
by: Cheng, Zhoujun, et al.
Published: (2026)
Scaling Optimal LR Across Token Horizons
by: Bjorck, Johan, et al.
Published: (2024)
by: Bjorck, Johan, et al.
Published: (2024)
Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
by: Huang, Audrey, et al.
Published: (2025)
by: Huang, Audrey, et al.
Published: (2025)
Compute Optimal Scaling of Skills: Knowledge vs Reasoning
by: Roberts, Nicholas, et al.
Published: (2025)
by: Roberts, Nicholas, et al.
Published: (2025)
Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
by: Abnar, Samira, et al.
Published: (2025)
by: Abnar, Samira, et al.
Published: (2025)
Direct Training Needs Regularisation: Anytime Optimal Inference Spiking Neural Network
by: Wu, Dengyu, et al.
Published: (2024)
by: Wu, Dengyu, et al.
Published: (2024)
Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
by: Li, Kevin Y., et al.
Published: (2024)
by: Li, Kevin Y., et al.
Published: (2024)
How Should LLMs Consume High-Quality Data? Optimal Data Scheduling via Quality-Aware Functional Scaling Laws
by: Zhu, Zhitao, et al.
Published: (2026)
by: Zhu, Zhitao, et al.
Published: (2026)
Disentangling Exploration of Large Language Models by Optimal Exploitation
by: Grams, Tim, et al.
Published: (2025)
by: Grams, Tim, et al.
Published: (2025)
Optimal Scaling Laws for Efficiency Gains in a Theoretical Transformer-Augmented Sectional MoE Framework
by: Sane, Soham
Published: (2025)
by: Sane, Soham
Published: (2025)
Training Free Guided Flow Matching with Optimal Control
by: Wang, Luran, et al.
Published: (2024)
by: Wang, Luran, et al.
Published: (2024)
Post-Norm can Resharpen Attention
by: Zsámboki, Pál, et al.
Published: (2025)
by: Zsámboki, Pál, et al.
Published: (2025)
Optimal Stability of KL Divergence under Gaussian Perturbations
by: Pan, Jialu, et al.
Published: (2026)
by: Pan, Jialu, et al.
Published: (2026)
Optimal Transport for LLM Reward Modeling from Noisy Preference
by: Pan, Licheng, et al.
Published: (2026)
by: Pan, Licheng, et al.
Published: (2026)
SOAR: Self-Correction for Optimal Alignment and Refinement in Diffusion Models
by: Qin, You, et al.
Published: (2026)
by: Qin, You, et al.
Published: (2026)
Efficient Controllable Diffusion via Optimal Classifier Guidance
by: Oertell, Owen, et al.
Published: (2025)
by: Oertell, Owen, et al.
Published: (2025)
BEACON: Bayesian Optimal Stopping for Efficient LLM Sampling
by: Wan, Guangya, et al.
Published: (2025)
by: Wan, Guangya, et al.
Published: (2025)
On Optimal Steering to Achieve Exact Fairness
by: Sharma, Mohit, et al.
Published: (2025)
by: Sharma, Mohit, et al.
Published: (2025)
Optimal patient allocation for echocardiographic assessments
by: Sun, Bozhi, et al.
Published: (2025)
by: Sun, Bozhi, et al.
Published: (2025)
Displacement-Sparse Neural Optimal Transport
by: Chen, Peter, et al.
Published: (2025)
by: Chen, Peter, et al.
Published: (2025)
Compute-Optimal Quantization-Aware Training
by: Dremov, Aleksandr, et al.
Published: (2025)
by: Dremov, Aleksandr, et al.
Published: (2025)
Optimal Policy Minimum Bayesian Risk
by: Astudillo, Ramón Fernandez, et al.
Published: (2025)
by: Astudillo, Ramón Fernandez, et al.
Published: (2025)
Distributional Counterfactual Explanations With Optimal Transport
by: You, Lei, et al.
Published: (2024)
by: You, Lei, et al.
Published: (2024)
GradientStabilizer:Fix the Norm, Not the Gradient
by: Huang, Tianjin, et al.
Published: (2025)
by: Huang, Tianjin, et al.
Published: (2025)
Benchmarking Reinforcement Learning via Stochastic Converse Optimality: Generating Systems with Known Optimal Policies
by: Ibrahim, Sinan, et al.
Published: (2026)
by: Ibrahim, Sinan, et al.
Published: (2026)
Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models
by: Fu, Yonggan, et al.
Published: (2025)
by: Fu, Yonggan, et al.
Published: (2025)
Solving Prior Distribution Mismatch in Diffusion Models via Optimal Transport
by: Wang, Zhanpeng, et al.
Published: (2024)
by: Wang, Zhanpeng, et al.
Published: (2024)
Training Neural Networks with Optimal Double-Bayesian Learning
by: Bui, Vy, et al.
Published: (2026)
by: Bui, Vy, et al.
Published: (2026)
Is Optimal Transport Necessary for Inverse Reinforcement Learning?
by: Dong, Zixuan, et al.
Published: (2025)
by: Dong, Zixuan, et al.
Published: (2025)
All AI Models are Wrong, but Some are Optimal
by: Anand, Akhil S, et al.
Published: (2025)
by: Anand, Akhil S, et al.
Published: (2025)
Similar Items
-
Time Transfer: On Optimal Learning Rate and Batch Size In The Infinite Data Limit
by: Filatov, Oleg, et al.
Published: (2024) -
Memory and Bandwidth are All You Need for Fully Sharded Data Parallel
by: Wang, Jiangtao, et al.
Published: (2025) -
OptScale: Probabilistic Optimality for Inference-time Scaling
by: Wang, Youkang, et al.
Published: (2025) -
Polynomial, trigonometric, and tropical activations
by: Khalfaoui-Hassani, Ismail, et al.
Published: (2025) -
MarkovScale: Towards Optimal Sequential Scaling at Inference Time
by: Wang, Youkang, et al.
Published: (2026)