Muon in Vision Transformers: Optimizer-Recipe Interactions and Gradient Spectra
Fuente:
arXiv
Salvato in:
| Autori principali: | Southworth, Ben S., Jiang, Shuai, McBride, Daniel, Cyr, Eric C., Thomas, Stephen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HoneyBee: Data Recipes for Vision-Language Reasoners
di: Bansal, Hritik, et al.
Pubblicazione: (2025)
di: Bansal, Hritik, et al.
Pubblicazione: (2025)
On the Convergence Behavior of Preconditioned Gradient Descent Toward the Rich Learning Regime
di: Jiang, Shuai, et al.
Pubblicazione: (2026)
di: Jiang, Shuai, et al.
Pubblicazione: (2026)
Beyond Muon: MUD (MomentUm Decorrelation) for Faster Transformer Training
di: Southworth, Ben S., et al.
Pubblicazione: (2026)
di: Southworth, Ben S., et al.
Pubblicazione: (2026)
On Vision Transformers for Classification Tasks in Side-Scan Sonar Imagery
di: Sheffield, BW, et al.
Pubblicazione: (2024)
di: Sheffield, BW, et al.
Pubblicazione: (2024)
RecipeGen: A Benchmark for Real-World Recipe Image Generation
di: Zhang, Ruoxuan, et al.
Pubblicazione: (2025)
di: Zhang, Ruoxuan, et al.
Pubblicazione: (2025)
SelaFD:Seamless Adaptation of Vision Transformer Fine-tuning for Radar-based Human Activity Recognition
di: Wang, Yijun, et al.
Pubblicazione: (2025)
di: Wang, Yijun, et al.
Pubblicazione: (2025)
RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation
di: Zhang, Ruoxuan, et al.
Pubblicazione: (2025)
di: Zhang, Ruoxuan, et al.
Pubblicazione: (2025)
PRUE: A Practical Recipe for Field Boundary Segmentation at Scale
di: Muhawenayo, Gedeon, et al.
Pubblicazione: (2026)
di: Muhawenayo, Gedeon, et al.
Pubblicazione: (2026)
Tarsier: Recipes for Training and Evaluating Large Video Description Models
di: Wang, Jiawei, et al.
Pubblicazione: (2024)
di: Wang, Jiawei, et al.
Pubblicazione: (2024)
Deep Image-to-Recipe Translation
di: Ma, Jiangqin, et al.
Pubblicazione: (2024)
di: Ma, Jiangqin, et al.
Pubblicazione: (2024)
HEAL-SWIN: A Vision Transformer On The Sphere
di: Carlsson, Oscar, et al.
Pubblicazione: (2023)
di: Carlsson, Oscar, et al.
Pubblicazione: (2023)
LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions
di: Mehri, Faridoun, et al.
Pubblicazione: (2024)
di: Mehri, Faridoun, et al.
Pubblicazione: (2024)
Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization
di: Jiang, Jiarui, et al.
Pubblicazione: (2024)
di: Jiang, Jiarui, et al.
Pubblicazione: (2024)
Efficient Training of Diffusion Mixture-of-Experts Models: A Practical Recipe
di: Liu, Yahui, et al.
Pubblicazione: (2025)
di: Liu, Yahui, et al.
Pubblicazione: (2025)
Mixture-of-Experts Models in Vision: Routing, Optimization, and Generalization
di: Rokah, Adam, et al.
Pubblicazione: (2026)
di: Rokah, Adam, et al.
Pubblicazione: (2026)
Native Segmentation Vision Transformers
di: Brasó, Guillem, et al.
Pubblicazione: (2025)
di: Brasó, Guillem, et al.
Pubblicazione: (2025)
A Deep Learning Approach to Estimate Canopy Height and Uncertainty by Integrating Seasonal Optical, SAR and Limited GEDI LiDAR Data over Northern Forests
di: Castro, Jose B., et al.
Pubblicazione: (2024)
di: Castro, Jose B., et al.
Pubblicazione: (2024)
A Recipe for Unbounded Data Augmentation in Visual Reinforcement Learning
di: Almuzairee, Abdulaziz, et al.
Pubblicazione: (2024)
di: Almuzairee, Abdulaziz, et al.
Pubblicazione: (2024)
Iwin Transformer: Hierarchical Vision Transformer using Interleaved Windows
di: Huo, Simin, et al.
Pubblicazione: (2025)
di: Huo, Simin, et al.
Pubblicazione: (2025)
Feedback Alignment Meets Low-Rank Manifolds: A Structured Recipe for Local Learning
di: Roy, Arani, et al.
Pubblicazione: (2025)
di: Roy, Arani, et al.
Pubblicazione: (2025)
Slicing Vision Transformer for Flexible Inference
di: Zhang, Yitian, et al.
Pubblicazione: (2024)
di: Zhang, Yitian, et al.
Pubblicazione: (2024)
RAViT: Resolution-Adaptive Vision Transformer
di: Guidez, Martial, et al.
Pubblicazione: (2026)
di: Guidez, Martial, et al.
Pubblicazione: (2026)
Rotary Position Embedding for Vision Transformer
di: Heo, Byeongho, et al.
Pubblicazione: (2024)
di: Heo, Byeongho, et al.
Pubblicazione: (2024)
Decorrelation Speeds Up Vision Transformers
di: Carrigg, Kieran, et al.
Pubblicazione: (2025)
di: Carrigg, Kieran, et al.
Pubblicazione: (2025)
Hybrid Convolution and Vision Transformer NAS Search Space for TinyML Image Classification
di: Djajapermana, Mikhael, et al.
Pubblicazione: (2025)
di: Djajapermana, Mikhael, et al.
Pubblicazione: (2025)
ZENITH: Automated Gradient Norm Informed Stochastic Optimization
di: Saha, Dhrubo
Pubblicazione: (2026)
di: Saha, Dhrubo
Pubblicazione: (2026)
Retrieval Augmented Recipe Generation
di: Liu, Guoshan, et al.
Pubblicazione: (2024)
di: Liu, Guoshan, et al.
Pubblicazione: (2024)
CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling
di: Wang, Xinze, et al.
Pubblicazione: (2025)
di: Wang, Xinze, et al.
Pubblicazione: (2025)
Block-Recurrent Dynamics in Vision Transformers
di: Jacobs, Mozes, et al.
Pubblicazione: (2025)
di: Jacobs, Mozes, et al.
Pubblicazione: (2025)
A Comparative Survey of Vision Transformers for Feature Extraction in Texture Analysis
di: Scabini, Leonardo, et al.
Pubblicazione: (2024)
di: Scabini, Leonardo, et al.
Pubblicazione: (2024)
SPoT: Subpixel Placement of Tokens in Vision Transformers
di: Hjelkrem-Tan, Martine, et al.
Pubblicazione: (2025)
di: Hjelkrem-Tan, Martine, et al.
Pubblicazione: (2025)
Instance-Aware Group Quantization for Vision Transformers
di: Moon, Jaehyeon, et al.
Pubblicazione: (2024)
di: Moon, Jaehyeon, et al.
Pubblicazione: (2024)
Vision Transformer-based Adversarial Domain Adaptation
di: Li, Yahan, et al.
Pubblicazione: (2024)
di: Li, Yahan, et al.
Pubblicazione: (2024)
Elastic Attention Cores for Scalable Vision Transformers
di: Song, Alan Z., et al.
Pubblicazione: (2026)
di: Song, Alan Z., et al.
Pubblicazione: (2026)
Compact Vision Transformer by Reduction of Kernel Complexity
di: Wang, Yancheng, et al.
Pubblicazione: (2025)
di: Wang, Yancheng, et al.
Pubblicazione: (2025)
Attention Transfer Is Not Universally Effective for Vision Transformers
di: Qin, Huaiyuan, et al.
Pubblicazione: (2026)
di: Qin, Huaiyuan, et al.
Pubblicazione: (2026)
Split Adaptation for Pre-trained Vision Transformers
di: Wang, Lixu, et al.
Pubblicazione: (2025)
di: Wang, Lixu, et al.
Pubblicazione: (2025)
DASViT: Differentiable Architecture Search for Vision Transformer
di: Wu, Pengjin, et al.
Pubblicazione: (2025)
di: Wu, Pengjin, et al.
Pubblicazione: (2025)
Stable Vision Concept Transformers for Medical Diagnosis
di: Hu, Lijie, et al.
Pubblicazione: (2025)
di: Hu, Lijie, et al.
Pubblicazione: (2025)
Self-Supervised Vision Transformers for Writer Retrieval
di: Raven, Tim, et al.
Pubblicazione: (2024)
di: Raven, Tim, et al.
Pubblicazione: (2024)
Documenti analoghi
-
HoneyBee: Data Recipes for Vision-Language Reasoners
di: Bansal, Hritik, et al.
Pubblicazione: (2025) -
On the Convergence Behavior of Preconditioned Gradient Descent Toward the Rich Learning Regime
di: Jiang, Shuai, et al.
Pubblicazione: (2026) -
Beyond Muon: MUD (MomentUm Decorrelation) for Faster Transformer Training
di: Southworth, Ben S., et al.
Pubblicazione: (2026) -
On Vision Transformers for Classification Tasks in Side-Scan Sonar Imagery
di: Sheffield, BW, et al.
Pubblicazione: (2024) -
RecipeGen: A Benchmark for Real-World Recipe Image Generation
di: Zhang, Ruoxuan, et al.
Pubblicazione: (2025)