Muon Optimizer Accelerates Grokking
Fuente:
arXiv
Saved in:
| Main Authors: | Tveit, Amund, Remseth, Bjørn, Skogvold, Arve |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Grokking Beyond the Euclidean Norm of Model Parameters
by: Notsawo, Pascal Jr Tikeng, et al.
Published: (2025)
by: Notsawo, Pascal Jr Tikeng, et al.
Published: (2025)
The Geometric Inductive Bias of Grokking: Bypassing Phase Transitions via Architectural Topology
by: Yıldırım, Alper
Published: (2026)
by: Yıldırım, Alper
Published: (2026)
Grokking in the Wild: Data Augmentation for Real-World Multi-Hop Reasoning with Transformers
by: Abramov, Roman, et al.
Published: (2025)
by: Abramov, Roman, et al.
Published: (2025)
FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft Learning
by: Zhang, Yizhou, et al.
Published: (2025)
by: Zhang, Yizhou, et al.
Published: (2025)
Spectral Compact Training: Pre-Training Large Language Models via Permanent Truncated SVD and Stiefel QR Retraction
by: Kohlberger, Björn Roman
Published: (2026)
by: Kohlberger, Björn Roman
Published: (2026)
Accelerating Transformer Inference and Training with 2:4 Activation Sparsity
by: Haziza, Daniel, et al.
Published: (2025)
by: Haziza, Daniel, et al.
Published: (2025)
OFMU: Optimization-Driven Framework for Machine Unlearning
by: Asif, Sadia, et al.
Published: (2025)
by: Asif, Sadia, et al.
Published: (2025)
Interpretability-Guided Bi-objective Optimization: Aligning Accuracy and Explainability
by: Fouladi, Kasra, et al.
Published: (2026)
by: Fouladi, Kasra, et al.
Published: (2026)
PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization
by: Ding, Ruiyi, et al.
Published: (2026)
by: Ding, Ruiyi, et al.
Published: (2026)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
Bounding Hallucinations: Information-Theoretic Guarantees for RAG Systems via Merlin-Arthur Protocols
by: Deiseroth, Björn, et al.
Published: (2025)
by: Deiseroth, Björn, et al.
Published: (2025)
Regret-Aware Policy Optimization: Environment-Level Memory for Replay Suppression under Delayed Harm
by: Hiremath, Prakul Sunil
Published: (2026)
by: Hiremath, Prakul Sunil
Published: (2026)
CopySpec: Accelerating LLMs with Speculative Copy-and-Paste Without Compromising Quality
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025)
xInv: Explainable Optimization of Inverse Problems
by: Memery, Sean, et al.
Published: (2025)
by: Memery, Sean, et al.
Published: (2025)
Optimizing Fantasy Sports Team Selection with Deep Reinforcement Learning
by: Bhattacharjee, Shamik, et al.
Published: (2024)
by: Bhattacharjee, Shamik, et al.
Published: (2024)
REVOLVE: Optimizing AI Systems by Tracking Response Evolution in Textual Optimization
by: Zhang, Peiyan, et al.
Published: (2024)
by: Zhang, Peiyan, et al.
Published: (2024)
Behavior Learning (BL): Learning Hierarchical Optimization Structures from Data
by: Ma, Zhenyao, et al.
Published: (2026)
by: Ma, Zhenyao, et al.
Published: (2026)
Balancing Efficiency and Effectiveness: An LLM-Infused Approach for Optimized CTR Prediction
by: Zhang, Guoxiao, et al.
Published: (2024)
by: Zhang, Guoxiao, et al.
Published: (2024)
Optimizing Robustness and Accuracy in Mixture of Experts: A Dual-Model Approach
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
QuIDE: Mastering the Quantized Intelligence Trade-off via Active Optimization
by: Jiang, Xiantao
Published: (2026)
by: Jiang, Xiantao
Published: (2026)
End-to-End Optimization of LLM-Driven Multi-Agent Search Systems via Heterogeneous-Group-Based Reinforcement Learning
by: Chen, Guanzhong, et al.
Published: (2025)
by: Chen, Guanzhong, et al.
Published: (2025)
Efficient Contextual Preferential Bayesian Optimization with Historical Examples
by: Khan, Farha A., et al.
Published: (2022)
by: Khan, Farha A., et al.
Published: (2022)
APP: Accelerated Path Patching with Task-Specific Pruning
by: Andersen, Frauke, et al.
Published: (2025)
by: Andersen, Frauke, et al.
Published: (2025)
DELTA: Variational Disentangled Learning for Privacy-Preserving Data Reprogramming
by: Malarkkan, Arun Vignesh, et al.
Published: (2025)
by: Malarkkan, Arun Vignesh, et al.
Published: (2025)
Solve it with EASE
by: Viktorin, Adam, et al.
Published: (2025)
by: Viktorin, Adam, et al.
Published: (2025)
TensorOpera Router: A Multi-Model Router for Efficient LLM Inference
by: Stripelis, Dimitris, et al.
Published: (2024)
by: Stripelis, Dimitris, et al.
Published: (2024)
Faster by Design: Interactive Aerodynamics via Neural Surrogates Trained on Expert-Validated CFD
by: Thumiger, Nicholas, et al.
Published: (2026)
by: Thumiger, Nicholas, et al.
Published: (2026)
CortexCompile: Harnessing Cortical-Inspired Architectures for Enhanced Multi-Agent NLP Code Synthesis
by: Ramachandran, Gautham, et al.
Published: (2024)
by: Ramachandran, Gautham, et al.
Published: (2024)
MACS: Multi-Agent Reinforcement Learning for Optimization of Crystal Structures
by: Zamaraeva, Elena, et al.
Published: (2025)
by: Zamaraeva, Elena, et al.
Published: (2025)
AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency
by: Höth, Max Henning, et al.
Published: (2026)
by: Höth, Max Henning, et al.
Published: (2026)
MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
by: Lan, Guangchen, et al.
Published: (2025)
by: Lan, Guangchen, et al.
Published: (2025)
Accelerating Model-Based Reinforcement Learning with State-Space World Models
by: Krinner, Maria, et al.
Published: (2025)
by: Krinner, Maria, et al.
Published: (2025)
Structured Prompt Optimization Meets Reinforcement Learning for Global and Local Interpretability over Complex Text
by: Zhou, Tianyang, et al.
Published: (2026)
by: Zhou, Tianyang, et al.
Published: (2026)
LABO: LLM-Accelerated Bayesian Optimization through Broad Exploration and Selective Experimentation
by: Chen, Zhuo, et al.
Published: (2026)
by: Chen, Zhuo, et al.
Published: (2026)
QiMeng-TensorOp: Automatically Generating High-Performance Tensor Operators with Hardware Primitives
by: Zhang, Xuzhi, et al.
Published: (2025)
by: Zhang, Xuzhi, et al.
Published: (2025)
Antibody Design and Optimization with Multi-scale Equivariant Graph Diffusion Models for Accurate Complex Antigen Binding
by: Chen, Jiameng, et al.
Published: (2025)
by: Chen, Jiameng, et al.
Published: (2025)
MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization
by: Tanjim, Md Mehrab, et al.
Published: (2026)
by: Tanjim, Md Mehrab, et al.
Published: (2026)
CoATA: Effective Co-Augmentation of Topology and Attribute for Graph Neural Networks
by: Liu, Tao, et al.
Published: (2025)
by: Liu, Tao, et al.
Published: (2025)
Efficient Hierarchical Contrastive Self-supervising Learning for Time Series Classification via Importance-aware Resolution Selection
by: Garcia, Kevin, et al.
Published: (2025)
by: Garcia, Kevin, et al.
Published: (2025)
Inferring Reward Machines and Transition Machines from Partially Observable Markov Decision Processes
by: Wu, Yuly, et al.
Published: (2025)
by: Wu, Yuly, et al.
Published: (2025)
Similar Items
-
Grokking Beyond the Euclidean Norm of Model Parameters
by: Notsawo, Pascal Jr Tikeng, et al.
Published: (2025) -
The Geometric Inductive Bias of Grokking: Bypassing Phase Transitions via Architectural Topology
by: Yıldırım, Alper
Published: (2026) -
Grokking in the Wild: Data Augmentation for Real-World Multi-Hop Reasoning with Transformers
by: Abramov, Roman, et al.
Published: (2025) -
FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft Learning
by: Zhang, Yizhou, et al.
Published: (2025) -
Spectral Compact Training: Pre-Training Large Language Models via Permanent Truncated SVD and Stiefel QR Retraction
by: Kohlberger, Björn Roman
Published: (2026)