Sign-In to the Lottery: Reparameterizing Sparse Training From Scratch
Fuente:
arXiv
Saved in:
| Main Authors: | Gadhikar, Advait, Jacobs, Tom, Zhou, Chao, Burkholz, Rebekka |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cyclic Sparse Training: Is it Enough?
by: Gadhikar, Advait, et al.
Published: (2024)
by: Gadhikar, Advait, et al.
Published: (2024)
Masks, Signs, And Learning Rate Rewinding
by: Gadhikar, Advait, et al.
Published: (2024)
by: Gadhikar, Advait, et al.
Published: (2024)
Pay Attention to Small Weights
by: Zhou, Chao, et al.
Published: (2025)
by: Zhou, Chao, et al.
Published: (2025)
Hyperbolic Aware Minimization: Implicit Bias for Sparsity
by: Jacobs, Tom, et al.
Published: (2025)
by: Jacobs, Tom, et al.
Published: (2025)
Never Saddle for Reparameterized Steepest Descent as Mirror Flow
by: Jacobs, Tom, et al.
Published: (2026)
by: Jacobs, Tom, et al.
Published: (2026)
Bridging Domains through Subspace-Aware Model Merging
by: Chaves, Levy, et al.
Published: (2026)
by: Chaves, Levy, et al.
Published: (2026)
HORST: Composing Optimizer Geometries for Sparse Transformer Training
by: Jacobs, Tom, et al.
Published: (2026)
by: Jacobs, Tom, et al.
Published: (2026)
Bayesian Lottery Ticket Hypothesis
by: Kuhn, Nicholas, et al.
Published: (2026)
by: Kuhn, Nicholas, et al.
Published: (2026)
Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias?
by: Jacobs, Tom, et al.
Published: (2025)
by: Jacobs, Tom, et al.
Published: (2025)
Mask in the Mirror: Implicit Sparsification
by: Jacobs, Tom, et al.
Published: (2024)
by: Jacobs, Tom, et al.
Published: (2024)
Reparameterized Multi-Resolution Convolutions for Long Sequence Modelling
by: Cunningham, Harry Jake, et al.
Published: (2024)
by: Cunningham, Harry Jake, et al.
Published: (2024)
RECAST: Reparameterized, Compact weight Adaptation for Sequential Tasks
by: Tasnim, Nazia, et al.
Published: (2024)
by: Tasnim, Nazia, et al.
Published: (2024)
Auto-Train-Once: Controller Network Guided Automatic Network Pruning from Scratch
by: Wu, Xidong, et al.
Published: (2024)
by: Wu, Xidong, et al.
Published: (2024)
Diffusion Model from Scratch
by: Zhen, Wang, et al.
Published: (2024)
by: Zhen, Wang, et al.
Published: (2024)
Reparameterized Tensor Ring Functional Decomposition for Multi-Dimensional Data Recovery
by: Xu, Yangyang, et al.
Published: (2026)
by: Xu, Yangyang, et al.
Published: (2026)
Sparse-to-Sparse Training of Diffusion Models
by: Oliveira, Inês Cardoso, et al.
Published: (2025)
by: Oliveira, Inês Cardoso, et al.
Published: (2025)
Training Neural Networks from Scratch with Parallel Low-Rank Adapters
by: Huh, Minyoung, et al.
Published: (2024)
by: Huh, Minyoung, et al.
Published: (2024)
Routing the Lottery: Adaptive Subnetworks for Heterogeneous Data
by: Stefanski, Grzegorz, et al.
Published: (2026)
by: Stefanski, Grzegorz, et al.
Published: (2026)
Winning the Lottery by Preserving Network Training Dynamics with Concrete Ticket Search
by: Arora, Tanay, et al.
Published: (2025)
by: Arora, Tanay, et al.
Published: (2025)
Dynamic Sparse Training with Structured Sparsity
by: Lasby, Mike, et al.
Published: (2023)
by: Lasby, Mike, et al.
Published: (2023)
Stretching Each Dollar: Diffusion Training from Scratch on a Micro-Budget
by: Sehwag, Vikash, et al.
Published: (2024)
by: Sehwag, Vikash, et al.
Published: (2024)
Generating Potent Poisons and Backdoors from Scratch with Guided Diffusion
by: Souri, Hossein, et al.
Published: (2024)
by: Souri, Hossein, et al.
Published: (2024)
AugLift: Depth-Aware Input Reparameterization Improves Domain Generalization in 2D-to-3D Pose Lifting
by: Warner, Nikolai, et al.
Published: (2025)
by: Warner, Nikolai, et al.
Published: (2025)
SuperTickets: Drawing Task-Agnostic Lottery Tickets from Supernets via Jointly Architecture Searching and Parameter Pruning
by: You, Haoran, et al.
Published: (2022)
by: You, Haoran, et al.
Published: (2022)
Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme
by: Ma, Yan, et al.
Published: (2025)
by: Ma, Yan, et al.
Published: (2025)
Beyond Words: AuralLLM and SignMST-C for Sign Language Production and Bidirectional Accessibility
by: Li, Yulong, et al.
Published: (2025)
by: Li, Yulong, et al.
Published: (2025)
LVSA: Training-Free Sparse Attention for Long Video Diffusion
by: Glorian, Gael, et al.
Published: (2026)
by: Glorian, Gael, et al.
Published: (2026)
TinyTrain: Resource-Aware Task-Adaptive Sparse Training of DNNs at the Data-Scarce Edge
by: Kwon, Young D., et al.
Published: (2023)
by: Kwon, Young D., et al.
Published: (2023)
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
by: Li, Zhiqi, et al.
Published: (2025)
by: Li, Zhiqi, et al.
Published: (2025)
LOTUS: Improving Transformer Efficiency with Sparsity Pruning and Data Lottery Tickets
by: Upadhyay, Ojasw
Published: (2024)
by: Upadhyay, Ojasw
Published: (2024)
Sparse-IFT: Sparse Iso-FLOP Transformations for Maximizing Training Efficiency
by: Thangarasa, Vithursan, et al.
Published: (2023)
by: Thangarasa, Vithursan, et al.
Published: (2023)
Attention Is All You Need For Mixture-of-Depths Routing
by: Gadhikar, Advait, et al.
Published: (2024)
by: Gadhikar, Advait, et al.
Published: (2024)
Early-Bird GCNs: Graph-Network Co-Optimization Towards More Efficient GCN Training and Inference via Drawing Early-Bird Lottery Tickets
by: You, Haoran, et al.
Published: (2021)
by: You, Haoran, et al.
Published: (2021)
Geo-Sign: Hyperbolic Contrastive Regularisation for Geometrically Aware Sign Language Translation
by: Fish, Edward, et al.
Published: (2025)
by: Fish, Edward, et al.
Published: (2025)
EfficientSign: An Attention-Enhanced Lightweight Architecture for Indian Sign Language Recognition
by: Gupta, Rishabh, et al.
Published: (2026)
by: Gupta, Rishabh, et al.
Published: (2026)
Scratching Visual Transformer's Back with Uniform Attention
by: Hyeon-Woo, Nam, et al.
Published: (2022)
by: Hyeon-Woo, Nam, et al.
Published: (2022)
Robust Experts: the Effect of Adversarial Training on CNNs with Sparse Mixture-of-Experts Layers
by: Pavlitska, Svetlana, et al.
Published: (2025)
by: Pavlitska, Svetlana, et al.
Published: (2025)
Embracing Unknown Step by Step: Towards Reliable Sparse Training in Real World
by: Lei, Bowen, et al.
Published: (2024)
by: Lei, Bowen, et al.
Published: (2024)
Does Vector Quantization Fail in Spatio-Temporal Forecasting? Exploring a Differentiable Sparse Soft-Vector Quantization Approach
by: Chen, Chao, et al.
Published: (2023)
by: Chen, Chao, et al.
Published: (2023)
A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)
by: Tu, Weijie, et al.
Published: (2024)
by: Tu, Weijie, et al.
Published: (2024)
Similar Items
-
Cyclic Sparse Training: Is it Enough?
by: Gadhikar, Advait, et al.
Published: (2024) -
Masks, Signs, And Learning Rate Rewinding
by: Gadhikar, Advait, et al.
Published: (2024) -
Pay Attention to Small Weights
by: Zhou, Chao, et al.
Published: (2025) -
Hyperbolic Aware Minimization: Implicit Bias for Sparsity
by: Jacobs, Tom, et al.
Published: (2025) -
Never Saddle for Reparameterized Steepest Descent as Mirror Flow
by: Jacobs, Tom, et al.
Published: (2026)