Hard ASH: Sparsity and the right optimizer make a continual learner
Fuente:
arXiv
Salvato in:
| Autore principale: | Keskinen, Santtu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Task agnostic continual learning with Pairwise layer architecture
di: Keskinen, Santtu
Pubblicazione: (2024)
di: Keskinen, Santtu
Pubblicazione: (2024)
Dynamic Sparse Training with Structured Sparsity
di: Lasby, Mike, et al.
Pubblicazione: (2023)
di: Lasby, Mike, et al.
Pubblicazione: (2023)
SINR: Sparsity Driven Compressed Implicit Neural Representations
di: Jayasundara, Dhananjaya, et al.
Pubblicazione: (2025)
di: Jayasundara, Dhananjaya, et al.
Pubblicazione: (2025)
LASERS: LAtent Space Encoding for Representations with Sparsity for Generative Modeling
di: Li, Xin, et al.
Pubblicazione: (2024)
di: Li, Xin, et al.
Pubblicazione: (2024)
Dual-Stage Invariant Continual Learning under Extreme Visual Sparsity
di: Zhang, Rangya, et al.
Pubblicazione: (2026)
di: Zhang, Rangya, et al.
Pubblicazione: (2026)
Rethinking Pruning for Vision-Language Models: Strategies for Effective Sparsity and Performance Restoration
di: He, Shwai, et al.
Pubblicazione: (2024)
di: He, Shwai, et al.
Pubblicazione: (2024)
OMH: Structured Sparsity via Optimally Matched Hierarchy for Unsupervised Semantic Segmentation
di: Ozaydin, Baran, et al.
Pubblicazione: (2024)
di: Ozaydin, Baran, et al.
Pubblicazione: (2024)
ELSA: Exploiting Layer-wise N:M Sparsity for Vision Transformer Acceleration
di: Huang, Ning-Chi, et al.
Pubblicazione: (2024)
di: Huang, Ning-Chi, et al.
Pubblicazione: (2024)
Sparsity Hurts: Simple Linear Adapter Can Boost Generalized Category Discovery
di: Ye, Bo, et al.
Pubblicazione: (2026)
di: Ye, Bo, et al.
Pubblicazione: (2026)
LAPA: Log-Domain Prediction-Driven Dynamic Sparsity Accelerator for Transformer Model
di: Wang, Huizheng, et al.
Pubblicazione: (2025)
di: Wang, Huizheng, et al.
Pubblicazione: (2025)
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
di: Xi, Haocheng, et al.
Pubblicazione: (2025)
di: Xi, Haocheng, et al.
Pubblicazione: (2025)
SURGEON: Memory-Adaptive Fully Test-Time Adaptation via Dynamic Activation Sparsity
di: Ma, Ke, et al.
Pubblicazione: (2025)
di: Ma, Ke, et al.
Pubblicazione: (2025)
Deep Ensembling with No Overhead for either Training or Testing: The All-Round Blessings of Dynamic Sparsity
di: Liu, Shiwei, et al.
Pubblicazione: (2021)
di: Liu, Shiwei, et al.
Pubblicazione: (2021)
VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
di: Wu, Zhenkai, et al.
Pubblicazione: (2025)
di: Wu, Zhenkai, et al.
Pubblicazione: (2025)
Meta-Sparsity: Learning Optimal Sparse Structures in Multi-task Networks through Meta-learning
di: Upadhyay, Richa, et al.
Pubblicazione: (2025)
di: Upadhyay, Richa, et al.
Pubblicazione: (2025)
Are Sparse Neural Networks Better Hard Sample Learners?
di: Xiao, Qiao, et al.
Pubblicazione: (2024)
di: Xiao, Qiao, et al.
Pubblicazione: (2024)
Improving Generalization via Meta-Learning on Hard Samples
di: Jain, Nishant, et al.
Pubblicazione: (2024)
di: Jain, Nishant, et al.
Pubblicazione: (2024)
RL makes MLLMs see better than SFT
di: Song, Junha, et al.
Pubblicazione: (2025)
di: Song, Junha, et al.
Pubblicazione: (2025)
FIS-DiT: Breaking the Few-Step Video Inference Barrier via Training-Free Frame Interleaved Sparsity
di: Tang, Jian, et al.
Pubblicazione: (2026)
di: Tang, Jian, et al.
Pubblicazione: (2026)
Rethinking Dataset Distillation: Hard Truths about Soft Labels
di: Dey, Priyam, et al.
Pubblicazione: (2026)
di: Dey, Priyam, et al.
Pubblicazione: (2026)
The Easy Path to Robustness: Coreset Selection using Sample Hardness
di: Ramesh, Pranav, et al.
Pubblicazione: (2025)
di: Ramesh, Pranav, et al.
Pubblicazione: (2025)
Hard Cases Detection in Motion Prediction by Vision-Language Foundation Models
di: Yang, Yi, et al.
Pubblicazione: (2024)
di: Yang, Yi, et al.
Pubblicazione: (2024)
HSMix: Hard and Soft Mixing Data Augmentation for Medical Image Segmentation
di: Sun, Danyang, et al.
Pubblicazione: (2025)
di: Sun, Danyang, et al.
Pubblicazione: (2025)
Video models are zero-shot learners and reasoners
di: Wiedemer, Thaddäus, et al.
Pubblicazione: (2025)
di: Wiedemer, Thaddäus, et al.
Pubblicazione: (2025)
Easy to Learn, Yet Hard to Forget: Towards Robust Unlearning Under Bias
di: Kwon, JuneHyoung, et al.
Pubblicazione: (2026)
di: Kwon, JuneHyoung, et al.
Pubblicazione: (2026)
Two out of Three (ToT): using self-consistency to make robust predictions
di: Lee, Jung Hoon, et al.
Pubblicazione: (2025)
di: Lee, Jung Hoon, et al.
Pubblicazione: (2025)
Extreme Model Compression with Structured Sparsity at Low Precision
di: Liu, Dan, et al.
Pubblicazione: (2025)
di: Liu, Dan, et al.
Pubblicazione: (2025)
Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying
di: Xue, Youze, et al.
Pubblicazione: (2025)
di: Xue, Youze, et al.
Pubblicazione: (2025)
Accelerating Targeted Hard-Label Adversarial Attacks in Low-Query Black-Box Settings
di: Swaminathan, Arjhun, et al.
Pubblicazione: (2025)
di: Swaminathan, Arjhun, et al.
Pubblicazione: (2025)
SHaSaM: Submodular Hard Sample Mining for Fair Facial Attribute Recognition
di: Majee, Anay, et al.
Pubblicazione: (2026)
di: Majee, Anay, et al.
Pubblicazione: (2026)
A New Approach for Evaluating and Improving the Performance of Segmentation Algorithms on Hard-to-Detect Blood Vessels
di: Parella, João Pedro, et al.
Pubblicazione: (2024)
di: Parella, João Pedro, et al.
Pubblicazione: (2024)
Active Learning for Finely-Categorized Image-Text Retrieval by Selecting Hard Negative Unpaired Samples
di: Jo, Dae Ung, et al.
Pubblicazione: (2024)
di: Jo, Dae Ung, et al.
Pubblicazione: (2024)
Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models
di: Huang, Xin, et al.
Pubblicazione: (2025)
di: Huang, Xin, et al.
Pubblicazione: (2025)
Adaptive Global and Fine-Grained Perceptual Fusion for MLLM Embeddings Compatible with Hard Negative Amplification
di: Hu, Lexiang, et al.
Pubblicazione: (2026)
di: Hu, Lexiang, et al.
Pubblicazione: (2026)
GD doesn't make the cut: Three ways that non-differentiability affects neural network training
di: Kumar, Siddharth Krishna
Pubblicazione: (2024)
di: Kumar, Siddharth Krishna
Pubblicazione: (2024)
From pre-training to downstream performance: Does domain-specific pre-training make sense?
di: Krones, Felix
Pubblicazione: (2026)
di: Krones, Felix
Pubblicazione: (2026)
CRISP: Hybrid Structured Sparsity for Class-aware Model Pruning
di: Aggarwal, Shivam, et al.
Pubblicazione: (2023)
di: Aggarwal, Shivam, et al.
Pubblicazione: (2023)
LOTUS: Improving Transformer Efficiency with Sparsity Pruning and Data Lottery Tickets
di: Upadhyay, Ojasw
Pubblicazione: (2024)
di: Upadhyay, Ojasw
Pubblicazione: (2024)
PLUM: Improving Inference Efficiency By Leveraging Repetition-Sparsity Trade-Off
di: Kuhar, Sachit, et al.
Pubblicazione: (2023)
di: Kuhar, Sachit, et al.
Pubblicazione: (2023)
Same accuracy, twice as fast: continuous training surpasses retraining from scratch
di: Verwimp, Eli, et al.
Pubblicazione: (2025)
di: Verwimp, Eli, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Task agnostic continual learning with Pairwise layer architecture
di: Keskinen, Santtu
Pubblicazione: (2024) -
Dynamic Sparse Training with Structured Sparsity
di: Lasby, Mike, et al.
Pubblicazione: (2023) -
SINR: Sparsity Driven Compressed Implicit Neural Representations
di: Jayasundara, Dhananjaya, et al.
Pubblicazione: (2025) -
LASERS: LAtent Space Encoding for Representations with Sparsity for Generative Modeling
di: Li, Xin, et al.
Pubblicazione: (2024) -
Dual-Stage Invariant Continual Learning under Extreme Visual Sparsity
di: Zhang, Rangya, et al.
Pubblicazione: (2026)