Saved in:
| Main Author: | Lo, Andy |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2411.01238 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Dynamic Structured Sparse Training with Learned Shuffles
by: Tyagi, Abhishek, et al.
Published: (2025)
by: Tyagi, Abhishek, et al.
Published: (2025)
Enhancing Transformer Training Efficiency with Dynamic Dropout
by: Yan, Hanrui, et al.
Published: (2024)
by: Yan, Hanrui, et al.
Published: (2024)
Grass: Compute Efficient Low-Memory LLM Training with Structured Sparse Gradients
by: Muhamed, Aashiq, et al.
Published: (2024)
by: Muhamed, Aashiq, et al.
Published: (2024)
DPad: Efficient Diffusion Language Models with Suffix Dropout
by: Chen, Xinhua, et al.
Published: (2025)
by: Chen, Xinhua, et al.
Published: (2025)
FairDropout: Using Example-Tied Dropout to Enhance Generalization of Minority Groups
by: Nanfack, Geraldin, et al.
Published: (2025)
by: Nanfack, Geraldin, et al.
Published: (2025)
Multi-level Monte Carlo Dropout for Efficient Uncertainty Quantification
by: Pim, Aaron, et al.
Published: (2026)
by: Pim, Aaron, et al.
Published: (2026)
Dropout-Based Rashomon Set Exploration for Efficient Predictive Multiplicity Estimation
by: Hsu, Hsiang, et al.
Published: (2024)
by: Hsu, Hsiang, et al.
Published: (2024)
Investigating the Synergistic Effects of Dropout and Residual Connections on Language Model Training
by: Li, Qingyang, et al.
Published: (2024)
by: Li, Qingyang, et al.
Published: (2024)
Progressive Data Dropout: An Embarrassingly Simple Approach to Faster Training
by: Sathiyanarayanan, Shriram M, et al.
Published: (2025)
by: Sathiyanarayanan, Shriram M, et al.
Published: (2025)
BiSparse-AAS: Bilinear Sparse Attention and Adaptive Spans Framework for Scalable and Efficient Text Summarization
by: Hagos, Desta Haileselassie, et al.
Published: (2025)
by: Hagos, Desta Haileselassie, et al.
Published: (2025)
Toward Efficient Influence Function: Dropout as a Compression Tool
by: Zhang, Yuchen, et al.
Published: (2025)
by: Zhang, Yuchen, et al.
Published: (2025)
Topology-Aware Revival for Efficient Sparse Training
by: Jin, Meiling, et al.
Published: (2026)
by: Jin, Meiling, et al.
Published: (2026)
Dropout Neural Network Training Viewed from a Percolation Perspective
by: Devlin, Finley, et al.
Published: (2025)
by: Devlin, Finley, et al.
Published: (2025)
Dynamic Sparse Training with Structured Sparsity
by: Lasby, Mike, et al.
Published: (2023)
by: Lasby, Mike, et al.
Published: (2023)
FedOBD: Opportunistic Block Dropout for Efficiently Training Large-scale Neural Networks through Federated Learning
by: Chen, Yuanyuan, et al.
Published: (2022)
by: Chen, Yuanyuan, et al.
Published: (2022)
Continuum Dropout for Neural Differential Equations
by: Lee, Jonghun, et al.
Published: (2025)
by: Lee, Jonghun, et al.
Published: (2025)
Dynamic Sparse Training of Diagonally Sparse Networks
by: Tyagi, Abhishek, et al.
Published: (2025)
by: Tyagi, Abhishek, et al.
Published: (2025)
Efficient Federated Learning with Heterogeneous Data and Adaptive Dropout
by: Liu, Ji, et al.
Published: (2025)
by: Liu, Ji, et al.
Published: (2025)
Dropout Drops Double Descent
by: Yang, Tian-Le, et al.
Published: (2023)
by: Yang, Tian-Le, et al.
Published: (2023)
Unreliable Uncertainty Estimates with Monte Carlo Dropout
by: Djupskås, Aslak, et al.
Published: (2025)
by: Djupskås, Aslak, et al.
Published: (2025)
Mixture-of-Channels: Exploiting Sparse FFNs for Efficient LLMs Pre-Training and Inference
by: Wu, Tong, et al.
Published: (2025)
by: Wu, Tong, et al.
Published: (2025)
ZO-SAM: Zero-Order Sharpness-Aware Minimization for Efficient Sparse Training
by: Ji, Jie, et al.
Published: (2026)
by: Ji, Jie, et al.
Published: (2026)
SPD-CFL: Stepwise Parameter Dropout for Efficient Continual Federated Learning
by: Yang, Yuning, et al.
Published: (2024)
by: Yang, Yuning, et al.
Published: (2024)
Localising Dropout Variance in Twin Networks
by: Doyle, Cooper
Published: (2025)
by: Doyle, Cooper
Published: (2025)
S2Aligner: Pair-Efficient and Transferable Pre-Training for Sparse Text-Attributed Graphs
by: Wang, Yuhan, et al.
Published: (2026)
by: Wang, Yuhan, et al.
Published: (2026)
Adaptive Tabu Dropout for Regularization of Deep Neural Network
by: Hasan, Md. Tarek, et al.
Published: (2024)
by: Hasan, Md. Tarek, et al.
Published: (2024)
An Adaptive Dropout Approach for High-Dimensional Bayesian Optimization
by: Huang, Jundi, et al.
Published: (2025)
by: Huang, Jundi, et al.
Published: (2025)
Dropout Induced Noise for Co-Creative GAN Systems
by: Wieluch, Sabine, et al.
Published: (2019)
by: Wieluch, Sabine, et al.
Published: (2019)
Generative Autoencoding of Dropout Patterns
by: Maeda, Shunta
Published: (2023)
by: Maeda, Shunta
Published: (2023)
SparseTransX: Efficient Training of Translation-Based Knowledge Graph Embeddings Using Sparse Matrix Operations
by: Anik, Md Saidul Hoque, et al.
Published: (2025)
by: Anik, Md Saidul Hoque, et al.
Published: (2025)
Sparse-ProxSkip: Accelerated Sparse-to-Sparse Training in Federated Learning
by: Meinhardt, Georg, et al.
Published: (2024)
by: Meinhardt, Georg, et al.
Published: (2024)
SparseOpt: Addressing Normalization-induced Gradient Skew in Sparse Training
by: Adnan, Mohammed, et al.
Published: (2026)
by: Adnan, Mohammed, et al.
Published: (2026)
Efficient Federated Fine-Tuning of Large Language Models with Layer Dropout
by: Wang, Shilong, et al.
Published: (2025)
by: Wang, Shilong, et al.
Published: (2025)
Federated Dropout: Convergence Analysis and Resource Allocation
by: Xie, Sijing, et al.
Published: (2024)
by: Xie, Sijing, et al.
Published: (2024)
Sparse-to-Sparse Training of Diffusion Models
by: Oliveira, Inês Cardoso, et al.
Published: (2025)
by: Oliveira, Inês Cardoso, et al.
Published: (2025)
Low Rank and Sparse Fourier Structure in Recurrent Networks Trained on Modular Addition
by: Rangamani, Akshay
Published: (2025)
by: Rangamani, Akshay
Published: (2025)
P-DROP: Poisson-Based Dropout for Graph Neural Networks
by: Yun, Hyunsik
Published: (2025)
by: Yun, Hyunsik
Published: (2025)
Effects of Dropout on Performance in Long-range Graph Learning Tasks
by: Singh, Jasraj, et al.
Published: (2025)
by: Singh, Jasraj, et al.
Published: (2025)
Understanding Transformer Encoder-Decoder Representations through Bernoulli Dropout
by: Chen, Xuanzhou
Published: (2026)
by: Chen, Xuanzhou
Published: (2026)
Fine-tuning with Very Large Dropout
by: Zhang, Jianyu, et al.
Published: (2024)
by: Zhang, Jianyu, et al.
Published: (2024)
Similar Items
-
Efficient Dynamic Structured Sparse Training with Learned Shuffles
by: Tyagi, Abhishek, et al.
Published: (2025) -
Enhancing Transformer Training Efficiency with Dynamic Dropout
by: Yan, Hanrui, et al.
Published: (2024) -
Grass: Compute Efficient Low-Memory LLM Training with Structured Sparse Gradients
by: Muhamed, Aashiq, et al.
Published: (2024) -
DPad: Efficient Diffusion Language Models with Suffix Dropout
by: Chen, Xinhua, et al.
Published: (2025) -
FairDropout: Using Example-Tied Dropout to Enhance Generalization of Minority Groups
by: Nanfack, Geraldin, et al.
Published: (2025)