Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Chongyu, Liu, Gaowen, Hong, Mingyi, Kompella, Ramana Rao, Liu, Sijia |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Trojan Horses to Castle Walls: Unveiling Bilateral Data Poisoning Effects in Diffusion Models
by: Pan, Zhuoshi, et al.
Published: (2023)
by: Pan, Zhuoshi, et al.
Published: (2023)
Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond
by: Fan, Chongyu, et al.
Published: (2025)
by: Fan, Chongyu, et al.
Published: (2025)
Towards Vector Optimization on Low-Dimensional Vector Symbolic Architecture
by: Duan, Shijin, et al.
Published: (2025)
by: Duan, Shijin, et al.
Published: (2025)
Riemannian Multinomial Logistics Regression for SPD Neural Networks
by: Chen, Ziheng, et al.
Published: (2023)
by: Chen, Ziheng, et al.
Published: (2023)
EPiC: Towards Lossless Speedup for Reasoning Training through Edge-Preserving CoT Condensation
by: Jia, Jinghan, et al.
Published: (2025)
by: Jia, Jinghan, et al.
Published: (2025)
Context Bootstrapped Reinforcement Learning
by: Agashe, Saaket, et al.
Published: (2026)
by: Agashe, Saaket, et al.
Published: (2026)
VLA Knows Its Limits
by: Wang, Haoxuan, et al.
Published: (2026)
by: Wang, Haoxuan, et al.
Published: (2026)
Vision Verification Enhanced Fusion of VLMs for Efficient Visual Reasoning
by: Tekin, Selim Furkan, et al.
Published: (2026)
by: Tekin, Selim Furkan, et al.
Published: (2026)
ProDiF: Protecting Domain-Invariant Features to Secure Pre-Trained Models Against Extraction
by: Zhou, Tong, et al.
Published: (2025)
by: Zhou, Tong, et al.
Published: (2025)
Subspace Control: Turning Constrained Model Steering into Controllable Spectral Optimization
by: Huang, Yancheng, et al.
Published: (2026)
by: Huang, Yancheng, et al.
Published: (2026)
ConQuER: Modular Architectures for Control and Bias Mitigation in IQP Quantum Generative Models
by: Zou, Xiaocheng, et al.
Published: (2025)
by: Zou, Xiaocheng, et al.
Published: (2025)
Sparse Mixture-of-Experts for Compositional Generalization: Empirical Evidence and Theoretical Foundations of Optimal Sparsity
by: Zhao, Jinze, et al.
Published: (2024)
by: Zhao, Jinze, et al.
Published: (2024)
Attention-aware Inference Optimizations for Large Vision-Language Models with Memory-efficient Decoding
by: Ilhan, Fatih, et al.
Published: (2026)
by: Ilhan, Fatih, et al.
Published: (2026)
Prompt Diffusion Robustifies Any-Modality Prompt Learning
by: Du, Yingjun, et al.
Published: (2024)
by: Du, Yingjun, et al.
Published: (2024)
Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
by: Fan, Chongyu, et al.
Published: (2024)
by: Fan, Chongyu, et al.
Published: (2024)
A First-order Generative Bilevel Optimization Framework for Diffusion Models
by: Xiao, Quan, et al.
Published: (2025)
by: Xiao, Quan, et al.
Published: (2025)
Challenging Forgets: Unveiling the Worst-Case Forget Sets in Machine Unlearning
by: Fan, Chongyu, et al.
Published: (2024)
by: Fan, Chongyu, et al.
Published: (2024)
MuonBP: Faster Muon via Block-Periodic Orthogonalization
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
by: Hao, Zhezheng, et al.
Published: (2025)
by: Hao, Zhezheng, et al.
Published: (2025)
LLM Unlearning Under the Microscope: A Full-Stack View on Methods and Metrics
by: Fan, Chongyu, et al.
Published: (2025)
by: Fan, Chongyu, et al.
Published: (2025)
On the Convergence of Muon and Beyond
by: Chang, Da, et al.
Published: (2025)
by: Chang, Da, et al.
Published: (2025)
Motion Marionette: Rethinking Rigid Motion Transfer via Prior Guidance
by: Wang, Haoxuan, et al.
Published: (2025)
by: Wang, Haoxuan, et al.
Published: (2025)
Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning
by: Lang, Yicheng, et al.
Published: (2025)
by: Lang, Yicheng, et al.
Published: (2025)
SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and Generation
by: Fan, Chongyu, et al.
Published: (2023)
by: Fan, Chongyu, et al.
Published: (2023)
Practical Efficiency of Muon for Pretraining
by: AI, Essential, et al.
Published: (2025)
by: AI, Essential, et al.
Published: (2025)
FedHFT: Efficient Federated Finetuning with Heterogeneous Edge Clients
by: Ilhan, Fatih, et al.
Published: (2025)
by: Ilhan, Fatih, et al.
Published: (2025)
UnlearnCanvas: Stylized Image Dataset for Enhanced Machine Unlearning Evaluation in Diffusion Models
by: Zhang, Yihua, et al.
Published: (2024)
by: Zhang, Yihua, et al.
Published: (2024)
Reversing the Forget-Retain Objectives: An Efficient LLM Unlearning Framework from Logit Difference
by: Ji, Jiabao, et al.
Published: (2024)
by: Ji, Jiabao, et al.
Published: (2024)
TEON: Tensorized Orthonormalization Beyond Layer-Wise Muon for Large Language Model Pre-Training
by: Zhang, Ruijie, et al.
Published: (2026)
by: Zhang, Ruijie, et al.
Published: (2026)
Pass@k Metric for RLVR: A Diagnostic Tool of Exploration, But Not an Objective
by: Yu, Yang
Published: (2025)
by: Yu, Yang
Published: (2025)
Not All Federated Learning Algorithms Are Created Equal: A Performance Evaluation Study
by: Baumgart, Gustav A., et al.
Published: (2024)
by: Baumgart, Gustav A., et al.
Published: (2024)
VL Norm: Rethink Loss Aggregation in RLVR
by: He, Zhiyuan, et al.
Published: (2025)
by: He, Zhiyuan, et al.
Published: (2025)
Spurious Rewards: Rethinking Training Signals in RLVR
by: Shao, Rulin, et al.
Published: (2025)
by: Shao, Rulin, et al.
Published: (2025)
Graphene: Infrastructure Security Posture Analysis with AI-generated Attack Graphs
by: Jin, Xin, et al.
Published: (2023)
by: Jin, Xin, et al.
Published: (2023)
Leak@$k$: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding
by: Reisizadeh, Hadi, et al.
Published: (2025)
by: Reisizadeh, Hadi, et al.
Published: (2025)
Variance-Adaptive Muon: Accelerating LLM Pretraining with NSR-Modulated and Variance-Scaled Momentum
by: Li, Jingru, et al.
Published: (2026)
by: Li, Jingru, et al.
Published: (2026)
Muon Optimizes Under Spectral Norm Constraints
by: Chen, Lizhang, et al.
Published: (2025)
by: Chen, Lizhang, et al.
Published: (2025)
Flame: Simplifying Topology Extension in Federated Learning
by: Daga, Harshit, et al.
Published: (2023)
by: Daga, Harshit, et al.
Published: (2023)
Efficient Multitask Dense Predictor via Binarization
by: Shang, Yuzhang, et al.
Published: (2024)
by: Shang, Yuzhang, et al.
Published: (2024)
Targeted Forgetting of Image Subgroups in CLIP Models
by: Zhang, Zeliang, et al.
Published: (2025)
by: Zhang, Zeliang, et al.
Published: (2025)
Similar Items
-
From Trojan Horses to Castle Walls: Unveiling Bilateral Data Poisoning Effects in Diffusion Models
by: Pan, Zhuoshi, et al.
Published: (2023) -
Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond
by: Fan, Chongyu, et al.
Published: (2025) -
Towards Vector Optimization on Low-Dimensional Vector Symbolic Architecture
by: Duan, Shijin, et al.
Published: (2025) -
Riemannian Multinomial Logistics Regression for SPD Neural Networks
by: Chen, Ziheng, et al.
Published: (2023) -
EPiC: Towards Lossless Speedup for Reasoning Training through Edge-Preserving CoT Condensation
by: Jia, Jinghan, et al.
Published: (2025)