How Controlling the Variance can Improve Training Stability of Sparsely Activated DNNs and CNNs
Fuente:
arXiv
Salvato in:
| Autori principali: | Dent, Emily, Tanner, Jared |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Theory of Minimal Weight Perturbations in Deep Networks and its Applications for Low-Rank Activated Backdoor Attacks
di: Evans, Bethan, et al.
Pubblicazione: (2026)
di: Evans, Bethan, et al.
Pubblicazione: (2026)
Model Merging by Output-Space Projection
di: Evans, Bethan, et al.
Pubblicazione: (2026)
di: Evans, Bethan, et al.
Pubblicazione: (2026)
Fixed-Confidence Best Arm Identification with Decreasing Variance
di: Roychowdhury, Tamojeet, et al.
Pubblicazione: (2025)
di: Roychowdhury, Tamojeet, et al.
Pubblicazione: (2025)
Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training
di: Wang, Yuanyi, et al.
Pubblicazione: (2026)
di: Wang, Yuanyi, et al.
Pubblicazione: (2026)
Sparse graphs using exchangeable random measures
di: Caron, François, et al.
Pubblicazione: (2014)
di: Caron, François, et al.
Pubblicazione: (2014)
The Price of Sparsity: Sufficient Conditions for Sparse Recovery using Sparse and Sparsified Measurements
di: Chaabouni, Youssef, et al.
Pubblicazione: (2025)
di: Chaabouni, Youssef, et al.
Pubblicazione: (2025)
Sparse Max-Affine Regression
di: Kanj, Haitham, et al.
Pubblicazione: (2024)
di: Kanj, Haitham, et al.
Pubblicazione: (2024)
Information Consistent Pruning: How to Efficiently Search for Sparse Networks?
di: Gharatappeh, Soheil, et al.
Pubblicazione: (2025)
di: Gharatappeh, Soheil, et al.
Pubblicazione: (2025)
A Unified Fractional Regularization Framework for Sparse Recovery
di: Zhao, Yinhao, et al.
Pubblicazione: (2026)
di: Zhao, Yinhao, et al.
Pubblicazione: (2026)
Optimal Variance-Dependent Regret Bounds for Infinite-Horizon MDPs
di: Zamir, Guy, et al.
Pubblicazione: (2026)
di: Zamir, Guy, et al.
Pubblicazione: (2026)
General Information Metrics for Improving AI Model Training Efficiency
di: Xu, Jianfeng, et al.
Pubblicazione: (2025)
di: Xu, Jianfeng, et al.
Pubblicazione: (2025)
Exact Recovery of Sparse Binary Vectors from Generalized Linear Measurements
di: Mazumdar, Arya, et al.
Pubblicazione: (2025)
di: Mazumdar, Arya, et al.
Pubblicazione: (2025)
A Unified Probabilistic Framework for Dictionary Learning with Parsimonious Activation
di: Zhao, Zihui, et al.
Pubblicazione: (2025)
di: Zhao, Zihui, et al.
Pubblicazione: (2025)
Sparse In-Network Learning via Shortest-Path Backpropagation and Finite-Rate Gating
di: Salehi, Mohammad Reza Deylam
Pubblicazione: (2026)
di: Salehi, Mohammad Reza Deylam
Pubblicazione: (2026)
LWM-Temporal: Sparse Spatio-Temporal Attention for Wireless Channel Representation Learning
di: Alikhani, Sadjad, et al.
Pubblicazione: (2026)
di: Alikhani, Sadjad, et al.
Pubblicazione: (2026)
Distributed Nonparametric Estimation: from Sparse to Dense Samples per Terminal
di: Yuan, Deheng, et al.
Pubblicazione: (2025)
di: Yuan, Deheng, et al.
Pubblicazione: (2025)
Transformers Provably Learn Sparse Token Selection While Fully-Connected Nets Cannot
di: Wang, Zixuan, et al.
Pubblicazione: (2024)
di: Wang, Zixuan, et al.
Pubblicazione: (2024)
Price of Quality: Sufficient Conditions for Sparse Recovery using Mixed-Quality Data
di: Chaabouni, Youssef, et al.
Pubblicazione: (2026)
di: Chaabouni, Youssef, et al.
Pubblicazione: (2026)
Competing Bandits in Matching Markets via Super Stability
di: Basu, Soumya
Pubblicazione: (2025)
di: Basu, Soumya
Pubblicazione: (2025)
Rate of Model Collapse in Recursive Training
di: Suresh, Ananda Theertha, et al.
Pubblicazione: (2024)
di: Suresh, Ananda Theertha, et al.
Pubblicazione: (2024)
Adaptivity can help exponentially for shadow tomography
di: Chen, Sitan, et al.
Pubblicazione: (2024)
di: Chen, Sitan, et al.
Pubblicazione: (2024)
Property Inheritance for Subtensors in Tensor Train Decompositions
di: Cai, HanQin, et al.
Pubblicazione: (2025)
di: Cai, HanQin, et al.
Pubblicazione: (2025)
How Patterns Dictate Learnability in Sequential Data
di: Morawski, Mario, et al.
Pubblicazione: (2025)
di: Morawski, Mario, et al.
Pubblicazione: (2025)
StateSMix: Online Lossless Compression via Mamba State Space Models and Sparse N-gram Context Mixing
di: Tacconelli, Roberto
Pubblicazione: (2026)
di: Tacconelli, Roberto
Pubblicazione: (2026)
Stabilization of Perturbed Loss Function: Differential Privacy without Gradient Noise
di: Habib, Salman, et al.
Pubblicazione: (2025)
di: Habib, Salman, et al.
Pubblicazione: (2025)
Training-Free Rate-Distortion-Perception Traversal With Diffusion
di: Wang, Yuhan, et al.
Pubblicazione: (2026)
di: Wang, Yuhan, et al.
Pubblicazione: (2026)
On the Training Convergence of Transformers for In-Context Classification of Gaussian Mixtures
di: Shen, Wei, et al.
Pubblicazione: (2024)
di: Shen, Wei, et al.
Pubblicazione: (2024)
MIST: Mutual Information Estimation Via Supervised Training
di: Gritsai, German, et al.
Pubblicazione: (2025)
di: Gritsai, German, et al.
Pubblicazione: (2025)
How Big Should a Wireless Foundation Model Be?
di: Cheng, Wei-Lun, et al.
Pubblicazione: (2026)
di: Cheng, Wei-Lun, et al.
Pubblicazione: (2026)
How Transformers Learn Causal Structure with Gradient Descent
di: Nichani, Eshaan, et al.
Pubblicazione: (2024)
di: Nichani, Eshaan, et al.
Pubblicazione: (2024)
Friendly Attacks to Improve Channel Coding Reliability
di: Kurmukova, Anastasiia, et al.
Pubblicazione: (2024)
di: Kurmukova, Anastasiia, et al.
Pubblicazione: (2024)
CSRv2: Unlocking Ultra-Sparse Embeddings
di: Guo, Lixuan, et al.
Pubblicazione: (2026)
di: Guo, Lixuan, et al.
Pubblicazione: (2026)
Improved Sample Complexity Bounds for Diffusion Model Training
di: Gupta, Shivam, et al.
Pubblicazione: (2023)
di: Gupta, Shivam, et al.
Pubblicazione: (2023)
Stabilizing Private LASSO under Heterogeneous Covariates via Anisotropic Objective Perturbation
di: Tanzawa, Haruka, et al.
Pubblicazione: (2026)
di: Tanzawa, Haruka, et al.
Pubblicazione: (2026)
Avoiding the Price of Adaptivity: Inference in Linear Contextual Bandits via Stability
di: Praharaj, Samya, et al.
Pubblicazione: (2025)
di: Praharaj, Samya, et al.
Pubblicazione: (2025)
Direction of Arrival Estimation with Sparse Subarrays
di: Leite, W., et al.
Pubblicazione: (2024)
di: Leite, W., et al.
Pubblicazione: (2024)
Leveraging Code Automorphisms for Improved Syndrome-Based Neural Decoding
di: Bidan, Raphaël Le, et al.
Pubblicazione: (2026)
di: Bidan, Raphaël Le, et al.
Pubblicazione: (2026)
Improved Regret Bounds for Linear Bandits with Heavy-Tailed Rewards
di: Tajdini, Artin, et al.
Pubblicazione: (2025)
di: Tajdini, Artin, et al.
Pubblicazione: (2025)
Improved Information Theoretic Generalization Bounds for Distributed and Federated Learning
di: Barnes, L. P., et al.
Pubblicazione: (2022)
di: Barnes, L. P., et al.
Pubblicazione: (2022)
Structured IB: Improving Information Bottleneck with Structured Feature Learning
di: Yang, Hanzhe, et al.
Pubblicazione: (2024)
di: Yang, Hanzhe, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Theory of Minimal Weight Perturbations in Deep Networks and its Applications for Low-Rank Activated Backdoor Attacks
di: Evans, Bethan, et al.
Pubblicazione: (2026) -
Model Merging by Output-Space Projection
di: Evans, Bethan, et al.
Pubblicazione: (2026) -
Fixed-Confidence Best Arm Identification with Decreasing Variance
di: Roychowdhury, Tamojeet, et al.
Pubblicazione: (2025) -
Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training
di: Wang, Yuanyi, et al.
Pubblicazione: (2026) -
Sparse graphs using exchangeable random measures
di: Caron, François, et al.
Pubblicazione: (2014)