Masks, Signs, And Learning Rate Rewinding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gadhikar, Advait, Burkholz, Rebekka |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sign-In to the Lottery: Reparameterizing Sparse Training From Scratch
von: Gadhikar, Advait, et al.
Veröffentlicht: (2025)
von: Gadhikar, Advait, et al.
Veröffentlicht: (2025)
Cyclic Sparse Training: Is it Enough?
von: Gadhikar, Advait, et al.
Veröffentlicht: (2024)
von: Gadhikar, Advait, et al.
Veröffentlicht: (2024)
Pay Attention to Small Weights
von: Zhou, Chao, et al.
Veröffentlicht: (2025)
von: Zhou, Chao, et al.
Veröffentlicht: (2025)
Hyperbolic Aware Minimization: Implicit Bias for Sparsity
von: Jacobs, Tom, et al.
Veröffentlicht: (2025)
von: Jacobs, Tom, et al.
Veröffentlicht: (2025)
Mask in the Mirror: Implicit Sparsification
von: Jacobs, Tom, et al.
Veröffentlicht: (2024)
von: Jacobs, Tom, et al.
Veröffentlicht: (2024)
GATE: How to Keep Out Intrusive Neighbors
von: Mustafa, Nimrah, et al.
Veröffentlicht: (2024)
von: Mustafa, Nimrah, et al.
Veröffentlicht: (2024)
OptRot: Mitigating Weight Outliers via Data-Free Rotations for Post-Training Quantization
von: Gadhikar, Advait, et al.
Veröffentlicht: (2025)
von: Gadhikar, Advait, et al.
Veröffentlicht: (2025)
Fixed Aggregation Features Can Rival GNNs
von: Rubio-Madrigal, Celia, et al.
Veröffentlicht: (2026)
von: Rubio-Madrigal, Celia, et al.
Veröffentlicht: (2026)
Robustness of Mixtures of Experts to Feature Noise
von: Sun, Dong, et al.
Veröffentlicht: (2026)
von: Sun, Dong, et al.
Veröffentlicht: (2026)
HORST: Composing Optimizer Geometries for Sparse Transformer Training
von: Jacobs, Tom, et al.
Veröffentlicht: (2026)
von: Jacobs, Tom, et al.
Veröffentlicht: (2026)
Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias?
von: Jacobs, Tom, et al.
Veröffentlicht: (2025)
von: Jacobs, Tom, et al.
Veröffentlicht: (2025)
Never Saddle for Reparameterized Steepest Descent as Mirror Flow
von: Jacobs, Tom, et al.
Veröffentlicht: (2026)
von: Jacobs, Tom, et al.
Veröffentlicht: (2026)
FedRewind: Rewinding Continual Model Exchange for Decentralized Federated Learning
von: Palazzo, Luca, et al.
Veröffentlicht: (2024)
von: Palazzo, Luca, et al.
Veröffentlicht: (2024)
Spectral Graph Pruning Against Over-Squashing and Over-Smoothing
von: Jamadandi, Adarsh, et al.
Veröffentlicht: (2024)
von: Jamadandi, Adarsh, et al.
Veröffentlicht: (2024)
GNNs Getting ComFy: Community and Feature Similarity Guided Rewiring
von: Rubio-Madrigal, Celia, et al.
Veröffentlicht: (2025)
von: Rubio-Madrigal, Celia, et al.
Veröffentlicht: (2025)
Multi-Agent Systems are Mixtures of Experts: Who Becomes an Influencer?
von: Bause, Franka, et al.
Veröffentlicht: (2026)
von: Bause, Franka, et al.
Veröffentlicht: (2026)
Pruning neural network models for gene regulatory dynamics using data and domain knowledge
von: Hossain, Intekhab, et al.
Veröffentlicht: (2024)
von: Hossain, Intekhab, et al.
Veröffentlicht: (2024)
When Shift Happens - Confounding Is to Blame
von: Reddy, Abbavaram Gowtham, et al.
Veröffentlicht: (2025)
von: Reddy, Abbavaram Gowtham, et al.
Veröffentlicht: (2025)
Frequency-Based Hyperparameter Selection in Games
von: Sanyal, Aniket, et al.
Veröffentlicht: (2026)
von: Sanyal, Aniket, et al.
Veröffentlicht: (2026)
Descend or Rewind? Stochastic Gradient Descent Unlearning
von: Mu, Siqiao, et al.
Veröffentlicht: (2025)
von: Mu, Siqiao, et al.
Veröffentlicht: (2025)
The Graphon Limit Hypothesis: Understanding Neural Network Pruning via Infinite Width Analysis
von: Pham, Hoang, et al.
Veröffentlicht: (2025)
von: Pham, Hoang, et al.
Veröffentlicht: (2025)
Rewind-to-Delete: Certified Machine Unlearning for Nonconvex Functions
von: Mu, Siqiao, et al.
Veröffentlicht: (2024)
von: Mu, Siqiao, et al.
Veröffentlicht: (2024)
Bridging Domains through Subspace-Aware Model Merging
von: Chaves, Levy, et al.
Veröffentlicht: (2026)
von: Chaves, Levy, et al.
Veröffentlicht: (2026)
SparseOpt: Addressing Normalization-induced Gradient Skew in Sparse Training
von: Adnan, Mohammed, et al.
Veröffentlicht: (2026)
von: Adnan, Mohammed, et al.
Veröffentlicht: (2026)
Can Prompts Rewind Time for LLMs? Evaluating the Effectiveness of Prompted Knowledge Cutoffs
von: Gao, Xin, et al.
Veröffentlicht: (2025)
von: Gao, Xin, et al.
Veröffentlicht: (2025)
Divergent Ensemble Networks: Enhancing Uncertainty Estimation with Shared Representations and Independent Branching
von: Kharbanda, Arnav, et al.
Veröffentlicht: (2024)
von: Kharbanda, Arnav, et al.
Veröffentlicht: (2024)
Sharp Convergence Rates for Masked Diffusion Models
von: Liang, Yuchen, et al.
Veröffentlicht: (2026)
von: Liang, Yuchen, et al.
Veröffentlicht: (2026)
GraSSRep: Graph-Based Self-Supervised Learning for Repeat Detection in Metagenomic Assembly
von: Azizpour, Ali, et al.
Veröffentlicht: (2024)
von: Azizpour, Ali, et al.
Veröffentlicht: (2024)
MuonAll: Muon Variant for Efficient Finetuning of Large Language Models
von: Page, Saurabh, et al.
Veröffentlicht: (2025)
von: Page, Saurabh, et al.
Veröffentlicht: (2025)
Unmask It! AI-Generated Product Review Detection in Dravidian Languages
von: De, Somsubhra, et al.
Veröffentlicht: (2025)
von: De, Somsubhra, et al.
Veröffentlicht: (2025)
Effective and Lightweight Representation Learning for Link Sign Prediction in Signed Bipartite Graphs
von: Gu, Gyeongmin, et al.
Veröffentlicht: (2024)
von: Gu, Gyeongmin, et al.
Veröffentlicht: (2024)
CAVACHON: a hierarchical variational autoencoder to integrate multi-modal single-cell data
von: Hsieh, Ping-Han, et al.
Veröffentlicht: (2024)
von: Hsieh, Ping-Han, et al.
Veröffentlicht: (2024)
Mask the Redundancy: Evolving Masking Representation Learning for Multivariate Time-Series Clustering
von: Tan, Zexi, et al.
Veröffentlicht: (2025)
von: Tan, Zexi, et al.
Veröffentlicht: (2025)
Stochastic-Sign SGD for Federated Learning with Theoretical Guarantees
von: Jin, Richeng, et al.
Veröffentlicht: (2020)
von: Jin, Richeng, et al.
Veröffentlicht: (2020)
Emergent Symbolic Structure in Health Foundation Models: Extraction, Alignment, and Cross-Modal Transfer
von: Katuwal, Gajendra, et al.
Veröffentlicht: (2026)
von: Katuwal, Gajendra, et al.
Veröffentlicht: (2026)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
von: Collins, Liam, et al.
Veröffentlicht: (2024)
von: Collins, Liam, et al.
Veröffentlicht: (2024)
Where to Mask: Structure-Guided Masking for Graph Masked Autoencoders
von: Liu, Chuang, et al.
Veröffentlicht: (2024)
von: Liu, Chuang, et al.
Veröffentlicht: (2024)
EdgeMask-DG*: Learning Domain-Invariant Graph Structures via Adversarial Edge Masking
von: Bhattacharya, Rishabh, et al.
Veröffentlicht: (2026)
von: Bhattacharya, Rishabh, et al.
Veröffentlicht: (2026)
Efficient Approximate Posterior Sampling with Annealed Langevin Monte Carlo
von: Parulekar, Advait, et al.
Veröffentlicht: (2025)
von: Parulekar, Advait, et al.
Veröffentlicht: (2025)
Gradient-Sign Masking for Task Vector Transport Across Pre-Trained Models
von: Rinaldi, Filippo, et al.
Veröffentlicht: (2025)
von: Rinaldi, Filippo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Sign-In to the Lottery: Reparameterizing Sparse Training From Scratch
von: Gadhikar, Advait, et al.
Veröffentlicht: (2025) -
Cyclic Sparse Training: Is it Enough?
von: Gadhikar, Advait, et al.
Veröffentlicht: (2024) -
Pay Attention to Small Weights
von: Zhou, Chao, et al.
Veröffentlicht: (2025) -
Hyperbolic Aware Minimization: Implicit Bias for Sparsity
von: Jacobs, Tom, et al.
Veröffentlicht: (2025) -
Mask in the Mirror: Implicit Sparsification
von: Jacobs, Tom, et al.
Veröffentlicht: (2024)