Efficient Dynamic Structured Sparse Training with Learned Shuffles
Fuente:
arXiv
Salvato in:
| Autori principali: | Tyagi, Abhishek, Iyer, Arjun, Young, Liam, Renninger, William H, Kanan, Christopher, Zhu, Yuhao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Dynamic Sparse Training of Diagonally Sparse Networks
di: Tyagi, Abhishek, et al.
Pubblicazione: (2025)
di: Tyagi, Abhishek, et al.
Pubblicazione: (2025)
Efficient Sparse Training with Structured Dropout
di: Lo, Andy
Pubblicazione: (2024)
di: Lo, Andy
Pubblicazione: (2024)
Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
di: Zhu, Linghao, et al.
Pubblicazione: (2025)
di: Zhu, Linghao, et al.
Pubblicazione: (2025)
GraVAC: Adaptive Compression for Communication-Efficient Distributed DL Training
di: Tyagi, Sahil, et al.
Pubblicazione: (2023)
di: Tyagi, Sahil, et al.
Pubblicazione: (2023)
Scalable Structure Learning for Sparse Context-Specific Systems
di: Rios, Felix Leopoldo, et al.
Pubblicazione: (2024)
di: Rios, Felix Leopoldo, et al.
Pubblicazione: (2024)
Overcoming the Stability Gap in Continual Learning
di: Harun, Md Yousuf, et al.
Pubblicazione: (2023)
di: Harun, Md Yousuf, et al.
Pubblicazione: (2023)
Controlling Neural Collapse Enhances Out-of-Distribution Detection and Transfer Learning
di: Harun, Md Yousuf, et al.
Pubblicazione: (2025)
di: Harun, Md Yousuf, et al.
Pubblicazione: (2025)
A Good Start Matters: Enhancing Continual Learning with Data-Driven Weight Initialization
di: Harun, Md Yousuf, et al.
Pubblicazione: (2025)
di: Harun, Md Yousuf, et al.
Pubblicazione: (2025)
To Shuffle or not to Shuffle: Auditing DP-SGD with Shuffling
di: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Pubblicazione: (2024)
di: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Pubblicazione: (2024)
Thales: Formulating and Estimating Architectural Vulnerability Factors for DNN Accelerators
di: Tyagi, Abhishek, et al.
Pubblicazione: (2022)
di: Tyagi, Abhishek, et al.
Pubblicazione: (2022)
OccamNets: Mitigating Dataset Bias by Favoring Simpler Hypotheses
di: Shrestha, Robik, et al.
Pubblicazione: (2022)
di: Shrestha, Robik, et al.
Pubblicazione: (2022)
Dynamic Sparse Training with Structured Sparsity
di: Lasby, Mike, et al.
Pubblicazione: (2023)
di: Lasby, Mike, et al.
Pubblicazione: (2023)
GRASP: A Rehearsal Policy for Efficient Online Continual Learning
di: Harun, Md Yousuf, et al.
Pubblicazione: (2023)
di: Harun, Md Yousuf, et al.
Pubblicazione: (2023)
Are Bias Mitigation Techniques for Deep Learning Effective?
di: Shrestha, Robik, et al.
Pubblicazione: (2021)
di: Shrestha, Robik, et al.
Pubblicazione: (2021)
Grass: Compute Efficient Low-Memory LLM Training with Structured Sparse Gradients
di: Muhamed, Aashiq, et al.
Pubblicazione: (2024)
di: Muhamed, Aashiq, et al.
Pubblicazione: (2024)
Group and Shuffle: Efficient Structured Orthogonal Parametrization
di: Gorbunov, Mikhail, et al.
Pubblicazione: (2024)
di: Gorbunov, Mikhail, et al.
Pubblicazione: (2024)
Joint Learning of Linear Dynamical Systems under Smoothness Constraints
di: Tyagi, Hemant
Pubblicazione: (2024)
di: Tyagi, Hemant
Pubblicazione: (2024)
WASH: Train your Ensemble with Communication-Efficient Weight Shuffling, then Average
di: Fournier, Louis, et al.
Pubblicazione: (2024)
di: Fournier, Louis, et al.
Pubblicazione: (2024)
Rethinking the Role of Dynamic Sparse Training for Scalable Deep Reinforcement Learning
di: Ma, Guozheng, et al.
Pubblicazione: (2025)
di: Ma, Guozheng, et al.
Pubblicazione: (2025)
DART-ing Through the Drift: Dynamic Tracing of Knowledge Neurons for Adaptive Inference-Time Pruning
di: Tyagi, Abhishek, et al.
Pubblicazione: (2026)
di: Tyagi, Abhishek, et al.
Pubblicazione: (2026)
Adjusted Shuffling SARAH: Advancing Complexity Analysis via Dynamic Gradient Weighting
di: Nguyen, Duc Toan, et al.
Pubblicazione: (2025)
di: Nguyen, Duc Toan, et al.
Pubblicazione: (2025)
Efficient Learning of Quantum States Prepared With Few Non-Clifford Gates
di: Grewal, Sabee, et al.
Pubblicazione: (2023)
di: Grewal, Sabee, et al.
Pubblicazione: (2023)
Value-Based Deep Multi-Agent Reinforcement Learning with Dynamic Sparse Training
di: Hu, Pihe, et al.
Pubblicazione: (2024)
di: Hu, Pihe, et al.
Pubblicazione: (2024)
Learning Page Order in Shuffled WOO Releases
di: Kahraman, Efe, et al.
Pubblicazione: (2026)
di: Kahraman, Efe, et al.
Pubblicazione: (2026)
Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training
di: Tyagi, Sahil, et al.
Pubblicazione: (2026)
di: Tyagi, Sahil, et al.
Pubblicazione: (2026)
OmniLearn: A Framework for Distributed Deep Learning over Heterogeneous Clusters
di: Tyagi, Sahil, et al.
Pubblicazione: (2025)
di: Tyagi, Sahil, et al.
Pubblicazione: (2025)
SMA: Submodular Modality Aligner For Data Efficient Multimodal Learning
di: Pham, Truong, et al.
Pubblicazione: (2026)
di: Pham, Truong, et al.
Pubblicazione: (2026)
Efficient Learning of Quantum States Prepared With Few Non-Clifford Gates II: Single-Copy Measurements
di: Grewal, Sabee, et al.
Pubblicazione: (2023)
di: Grewal, Sabee, et al.
Pubblicazione: (2023)
Nesterov acceleration in benignly non-convex landscapes
di: Gupta, Kanan, et al.
Pubblicazione: (2024)
di: Gupta, Kanan, et al.
Pubblicazione: (2024)
Understanding the Effect of Noise in LLM Training Data with Algorithmic Chains of Thought
di: Havrilla, Alex, et al.
Pubblicazione: (2024)
di: Havrilla, Alex, et al.
Pubblicazione: (2024)
Can Kans (re)discover predictive models for Direct-Drive Laser Fusion?
di: Ejaz, Rahman, et al.
Pubblicazione: (2024)
di: Ejaz, Rahman, et al.
Pubblicazione: (2024)
Shuffling Momentum Gradient Algorithm for Convex Optimization
di: Tran, Trang H., et al.
Pubblicazione: (2024)
di: Tran, Trang H., et al.
Pubblicazione: (2024)
Topology-Aware Revival for Efficient Sparse Training
di: Jin, Meiling, et al.
Pubblicazione: (2026)
di: Jin, Meiling, et al.
Pubblicazione: (2026)
Temporal Patch Shuffle (TPS): Leveraging Patch-Level Shuffling to Boost Generalization and Robustness in Time Series Forecasting
di: Bakhshaliyev, Jafar, et al.
Pubblicazione: (2026)
di: Bakhshaliyev, Jafar, et al.
Pubblicazione: (2026)
Sparse-ProxSkip: Accelerated Sparse-to-Sparse Training in Federated Learning
di: Meinhardt, Georg, et al.
Pubblicazione: (2024)
di: Meinhardt, Georg, et al.
Pubblicazione: (2024)
Dynamic Sparse Learning: A Novel Paradigm for Efficient Recommendation
di: Wang, Shuyao, et al.
Pubblicazione: (2024)
di: Wang, Shuyao, et al.
Pubblicazione: (2024)
Dynamics of Transient Structure in In-Context Linear Regression Transformers
di: Carroll, Liam, et al.
Pubblicazione: (2025)
di: Carroll, Liam, et al.
Pubblicazione: (2025)
Disentangling Dense Embeddings with Sparse Autoencoders
di: O'Neill, Charles, et al.
Pubblicazione: (2024)
di: O'Neill, Charles, et al.
Pubblicazione: (2024)
Sparse Bayesian Deep Functional Learning with Structured Region Selection
di: Zhu, Xiaoxian, et al.
Pubblicazione: (2026)
di: Zhu, Xiaoxian, et al.
Pubblicazione: (2026)
Improving Multimodal Large Language Models Using Continual Learning
di: Srivastava, Shikhar, et al.
Pubblicazione: (2024)
di: Srivastava, Shikhar, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Dynamic Sparse Training of Diagonally Sparse Networks
di: Tyagi, Abhishek, et al.
Pubblicazione: (2025) -
Efficient Sparse Training with Structured Dropout
di: Lo, Andy
Pubblicazione: (2024) -
Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
di: Zhu, Linghao, et al.
Pubblicazione: (2025) -
GraVAC: Adaptive Compression for Communication-Efficient Distributed DL Training
di: Tyagi, Sahil, et al.
Pubblicazione: (2023) -
Scalable Structure Learning for Sparse Context-Specific Systems
di: Rios, Felix Leopoldo, et al.
Pubblicazione: (2024)