HORST: Composing Optimizer Geometries for Sparse Transformer Training
Fuente:
arXiv
Saved in:
| Main Authors: | Jacobs, Tom, Jain, Rohan, Burkholz, Rebekka |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SparseOpt: Addressing Normalization-induced Gradient Skew in Sparse Training
by: Adnan, Mohammed, et al.
Published: (2026)
by: Adnan, Mohammed, et al.
Published: (2026)
Mask in the Mirror: Implicit Sparsification
by: Jacobs, Tom, et al.
Published: (2024)
by: Jacobs, Tom, et al.
Published: (2024)
Sign-In to the Lottery: Reparameterizing Sparse Training From Scratch
by: Gadhikar, Advait, et al.
Published: (2025)
by: Gadhikar, Advait, et al.
Published: (2025)
Never Saddle for Reparameterized Steepest Descent as Mirror Flow
by: Jacobs, Tom, et al.
Published: (2026)
by: Jacobs, Tom, et al.
Published: (2026)
Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias?
by: Jacobs, Tom, et al.
Published: (2025)
by: Jacobs, Tom, et al.
Published: (2025)
Cyclic Sparse Training: Is it Enough?
by: Gadhikar, Advait, et al.
Published: (2024)
by: Gadhikar, Advait, et al.
Published: (2024)
Pay Attention to Small Weights
by: Zhou, Chao, et al.
Published: (2025)
by: Zhou, Chao, et al.
Published: (2025)
Hyperbolic Aware Minimization: Implicit Bias for Sparsity
by: Jacobs, Tom, et al.
Published: (2025)
by: Jacobs, Tom, et al.
Published: (2025)
GATE: How to Keep Out Intrusive Neighbors
by: Mustafa, Nimrah, et al.
Published: (2024)
by: Mustafa, Nimrah, et al.
Published: (2024)
Masks, Signs, And Learning Rate Rewinding
by: Gadhikar, Advait, et al.
Published: (2024)
by: Gadhikar, Advait, et al.
Published: (2024)
The Graphon Limit Hypothesis: Understanding Neural Network Pruning via Infinite Width Analysis
by: Pham, Hoang, et al.
Published: (2025)
by: Pham, Hoang, et al.
Published: (2025)
Fixed Aggregation Features Can Rival GNNs
by: Rubio-Madrigal, Celia, et al.
Published: (2026)
by: Rubio-Madrigal, Celia, et al.
Published: (2026)
Robustness of Mixtures of Experts to Feature Noise
by: Sun, Dong, et al.
Published: (2026)
by: Sun, Dong, et al.
Published: (2026)
Spectral Graph Pruning Against Over-Squashing and Over-Smoothing
by: Jamadandi, Adarsh, et al.
Published: (2024)
by: Jamadandi, Adarsh, et al.
Published: (2024)
GNNs Getting ComFy: Community and Feature Similarity Guided Rewiring
by: Rubio-Madrigal, Celia, et al.
Published: (2025)
by: Rubio-Madrigal, Celia, et al.
Published: (2025)
Multi-Agent Systems are Mixtures of Experts: Who Becomes an Influencer?
by: Bause, Franka, et al.
Published: (2026)
by: Bause, Franka, et al.
Published: (2026)
Pruning neural network models for gene regulatory dynamics using data and domain knowledge
by: Hossain, Intekhab, et al.
Published: (2024)
by: Hossain, Intekhab, et al.
Published: (2024)
Frequency-Based Hyperparameter Selection in Games
by: Sanyal, Aniket, et al.
Published: (2026)
by: Sanyal, Aniket, et al.
Published: (2026)
When Shift Happens - Confounding Is to Blame
by: Reddy, Abbavaram Gowtham, et al.
Published: (2025)
by: Reddy, Abbavaram Gowtham, et al.
Published: (2025)
Implicit Bias of Mirror Flow in Homogeneous Neural Networks: Sparse and Dense Feature Learning
by: Jacobs, Tom, et al.
Published: (2026)
by: Jacobs, Tom, et al.
Published: (2026)
Sparse Training from Random Initialization: Aligning Lottery Ticket Masks using Weight Symmetry
by: Adnan, Mohammed, et al.
Published: (2025)
by: Adnan, Mohammed, et al.
Published: (2025)
Bridging Domains through Subspace-Aware Model Merging
by: Chaves, Levy, et al.
Published: (2026)
by: Chaves, Levy, et al.
Published: (2026)
PBSCR: The Piano Bootleg Score Composer Recognition Dataset
by: Jain, Arhan, et al.
Published: (2024)
by: Jain, Arhan, et al.
Published: (2024)
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
by: Panda, Ashwinee, et al.
Published: (2025)
by: Panda, Ashwinee, et al.
Published: (2025)
Dynamic Sparse Training of Diagonally Sparse Networks
by: Tyagi, Abhishek, et al.
Published: (2025)
by: Tyagi, Abhishek, et al.
Published: (2025)
Improving Transformers with Dynamically Composable Multi-Head Attention
by: Xiao, Da, et al.
Published: (2024)
by: Xiao, Da, et al.
Published: (2024)
Second-Order, First-Class: A Composable Stack for Curvature-Aware Training
by: Korbit, Mikalai, et al.
Published: (2026)
by: Korbit, Mikalai, et al.
Published: (2026)
On the Learn-to-Optimize Capabilities of Transformers in In-Context Sparse Recovery
by: Liu, Renpu, et al.
Published: (2024)
by: Liu, Renpu, et al.
Published: (2024)
Multi-Phase Spacecraft Trajectory Optimization via Transformer-Based Reinforcement Learning
by: Jain, Amit, et al.
Published: (2025)
by: Jain, Amit, et al.
Published: (2025)
ConText-CIR: Learning from Concepts in Text for Composed Image Retrieval
by: Xing, Eric, et al.
Published: (2025)
by: Xing, Eric, et al.
Published: (2025)
A Comparative Analysis of Transformer Models in Social Bot Detection
by: Veit, Rohan, et al.
Published: (2025)
by: Veit, Rohan, et al.
Published: (2025)
Advancing Software Engineering in the AI-ML Paradigm: A Study of Optimized Methodologies and Obstacles
by: Dr. Rohan Jain and Dr. Leela Rao
Published: (2023)
by: Dr. Rohan Jain and Dr. Leela Rao
Published: (2023)
Sparse-IFT: Sparse Iso-FLOP Transformations for Maximizing Training Efficiency
by: Thangarasa, Vithursan, et al.
Published: (2023)
by: Thangarasa, Vithursan, et al.
Published: (2023)
Extending Multi-Source Bayesian Optimization With Causality Principles
by: Jacobs, Luuk, et al.
Published: (2026)
by: Jacobs, Luuk, et al.
Published: (2026)
Efficient Resource-Constrained Training of Transformers via Subspace Optimization
by: Nguyen, Le-Trung, et al.
Published: (2025)
by: Nguyen, Le-Trung, et al.
Published: (2025)
Limits of Transformer Language Models on Learning to Compose Algorithms
by: Thomm, Jonathan, et al.
Published: (2024)
by: Thomm, Jonathan, et al.
Published: (2024)
SCALAR: Learning and Composing Skills through LLM Guided Symbolic Planning and Deep RL Grounding
by: Zabounidis, Renos, et al.
Published: (2026)
by: Zabounidis, Renos, et al.
Published: (2026)
Composer 2 Technical Report
by: Research, Cursor, et al.
Published: (2026)
by: Research, Cursor, et al.
Published: (2026)
Efficient Sparse Training with Structured Dropout
by: Lo, Andy
Published: (2024)
by: Lo, Andy
Published: (2024)
Understanding the Staged Dynamics of Transformers in Learning Latent Structure
by: Saha, Rohan, et al.
Published: (2025)
by: Saha, Rohan, et al.
Published: (2025)
Similar Items
-
SparseOpt: Addressing Normalization-induced Gradient Skew in Sparse Training
by: Adnan, Mohammed, et al.
Published: (2026) -
Mask in the Mirror: Implicit Sparsification
by: Jacobs, Tom, et al.
Published: (2024) -
Sign-In to the Lottery: Reparameterizing Sparse Training From Scratch
by: Gadhikar, Advait, et al.
Published: (2025) -
Never Saddle for Reparameterized Steepest Descent as Mirror Flow
by: Jacobs, Tom, et al.
Published: (2026) -
Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias?
by: Jacobs, Tom, et al.
Published: (2025)