Hard ASH: Sparsity and the right optimizer make a continual learner
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Keskinen, Santtu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Task agnostic continual learning with Pairwise layer architecture
von: Keskinen, Santtu
Veröffentlicht: (2024)
von: Keskinen, Santtu
Veröffentlicht: (2024)
Dynamic Sparse Training with Structured Sparsity
von: Lasby, Mike, et al.
Veröffentlicht: (2023)
von: Lasby, Mike, et al.
Veröffentlicht: (2023)
SINR: Sparsity Driven Compressed Implicit Neural Representations
von: Jayasundara, Dhananjaya, et al.
Veröffentlicht: (2025)
von: Jayasundara, Dhananjaya, et al.
Veröffentlicht: (2025)
LASERS: LAtent Space Encoding for Representations with Sparsity for Generative Modeling
von: Li, Xin, et al.
Veröffentlicht: (2024)
von: Li, Xin, et al.
Veröffentlicht: (2024)
Dual-Stage Invariant Continual Learning under Extreme Visual Sparsity
von: Zhang, Rangya, et al.
Veröffentlicht: (2026)
von: Zhang, Rangya, et al.
Veröffentlicht: (2026)
Rethinking Pruning for Vision-Language Models: Strategies for Effective Sparsity and Performance Restoration
von: He, Shwai, et al.
Veröffentlicht: (2024)
von: He, Shwai, et al.
Veröffentlicht: (2024)
OMH: Structured Sparsity via Optimally Matched Hierarchy for Unsupervised Semantic Segmentation
von: Ozaydin, Baran, et al.
Veröffentlicht: (2024)
von: Ozaydin, Baran, et al.
Veröffentlicht: (2024)
ELSA: Exploiting Layer-wise N:M Sparsity for Vision Transformer Acceleration
von: Huang, Ning-Chi, et al.
Veröffentlicht: (2024)
von: Huang, Ning-Chi, et al.
Veröffentlicht: (2024)
Sparsity Hurts: Simple Linear Adapter Can Boost Generalized Category Discovery
von: Ye, Bo, et al.
Veröffentlicht: (2026)
von: Ye, Bo, et al.
Veröffentlicht: (2026)
LAPA: Log-Domain Prediction-Driven Dynamic Sparsity Accelerator for Transformer Model
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
von: Xi, Haocheng, et al.
Veröffentlicht: (2025)
von: Xi, Haocheng, et al.
Veröffentlicht: (2025)
SURGEON: Memory-Adaptive Fully Test-Time Adaptation via Dynamic Activation Sparsity
von: Ma, Ke, et al.
Veröffentlicht: (2025)
von: Ma, Ke, et al.
Veröffentlicht: (2025)
Deep Ensembling with No Overhead for either Training or Testing: The All-Round Blessings of Dynamic Sparsity
von: Liu, Shiwei, et al.
Veröffentlicht: (2021)
von: Liu, Shiwei, et al.
Veröffentlicht: (2021)
VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
von: Wu, Zhenkai, et al.
Veröffentlicht: (2025)
von: Wu, Zhenkai, et al.
Veröffentlicht: (2025)
Meta-Sparsity: Learning Optimal Sparse Structures in Multi-task Networks through Meta-learning
von: Upadhyay, Richa, et al.
Veröffentlicht: (2025)
von: Upadhyay, Richa, et al.
Veröffentlicht: (2025)
Are Sparse Neural Networks Better Hard Sample Learners?
von: Xiao, Qiao, et al.
Veröffentlicht: (2024)
von: Xiao, Qiao, et al.
Veröffentlicht: (2024)
Improving Generalization via Meta-Learning on Hard Samples
von: Jain, Nishant, et al.
Veröffentlicht: (2024)
von: Jain, Nishant, et al.
Veröffentlicht: (2024)
RL makes MLLMs see better than SFT
von: Song, Junha, et al.
Veröffentlicht: (2025)
von: Song, Junha, et al.
Veröffentlicht: (2025)
FIS-DiT: Breaking the Few-Step Video Inference Barrier via Training-Free Frame Interleaved Sparsity
von: Tang, Jian, et al.
Veröffentlicht: (2026)
von: Tang, Jian, et al.
Veröffentlicht: (2026)
Rethinking Dataset Distillation: Hard Truths about Soft Labels
von: Dey, Priyam, et al.
Veröffentlicht: (2026)
von: Dey, Priyam, et al.
Veröffentlicht: (2026)
The Easy Path to Robustness: Coreset Selection using Sample Hardness
von: Ramesh, Pranav, et al.
Veröffentlicht: (2025)
von: Ramesh, Pranav, et al.
Veröffentlicht: (2025)
Hard Cases Detection in Motion Prediction by Vision-Language Foundation Models
von: Yang, Yi, et al.
Veröffentlicht: (2024)
von: Yang, Yi, et al.
Veröffentlicht: (2024)
HSMix: Hard and Soft Mixing Data Augmentation for Medical Image Segmentation
von: Sun, Danyang, et al.
Veröffentlicht: (2025)
von: Sun, Danyang, et al.
Veröffentlicht: (2025)
Video models are zero-shot learners and reasoners
von: Wiedemer, Thaddäus, et al.
Veröffentlicht: (2025)
von: Wiedemer, Thaddäus, et al.
Veröffentlicht: (2025)
Easy to Learn, Yet Hard to Forget: Towards Robust Unlearning Under Bias
von: Kwon, JuneHyoung, et al.
Veröffentlicht: (2026)
von: Kwon, JuneHyoung, et al.
Veröffentlicht: (2026)
Two out of Three (ToT): using self-consistency to make robust predictions
von: Lee, Jung Hoon, et al.
Veröffentlicht: (2025)
von: Lee, Jung Hoon, et al.
Veröffentlicht: (2025)
Extreme Model Compression with Structured Sparsity at Low Precision
von: Liu, Dan, et al.
Veröffentlicht: (2025)
von: Liu, Dan, et al.
Veröffentlicht: (2025)
Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying
von: Xue, Youze, et al.
Veröffentlicht: (2025)
von: Xue, Youze, et al.
Veröffentlicht: (2025)
Accelerating Targeted Hard-Label Adversarial Attacks in Low-Query Black-Box Settings
von: Swaminathan, Arjhun, et al.
Veröffentlicht: (2025)
von: Swaminathan, Arjhun, et al.
Veröffentlicht: (2025)
SHaSaM: Submodular Hard Sample Mining for Fair Facial Attribute Recognition
von: Majee, Anay, et al.
Veröffentlicht: (2026)
von: Majee, Anay, et al.
Veröffentlicht: (2026)
A New Approach for Evaluating and Improving the Performance of Segmentation Algorithms on Hard-to-Detect Blood Vessels
von: Parella, João Pedro, et al.
Veröffentlicht: (2024)
von: Parella, João Pedro, et al.
Veröffentlicht: (2024)
Active Learning for Finely-Categorized Image-Text Retrieval by Selecting Hard Negative Unpaired Samples
von: Jo, Dae Ung, et al.
Veröffentlicht: (2024)
von: Jo, Dae Ung, et al.
Veröffentlicht: (2024)
Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models
von: Huang, Xin, et al.
Veröffentlicht: (2025)
von: Huang, Xin, et al.
Veröffentlicht: (2025)
Adaptive Global and Fine-Grained Perceptual Fusion for MLLM Embeddings Compatible with Hard Negative Amplification
von: Hu, Lexiang, et al.
Veröffentlicht: (2026)
von: Hu, Lexiang, et al.
Veröffentlicht: (2026)
GD doesn't make the cut: Three ways that non-differentiability affects neural network training
von: Kumar, Siddharth Krishna
Veröffentlicht: (2024)
von: Kumar, Siddharth Krishna
Veröffentlicht: (2024)
From pre-training to downstream performance: Does domain-specific pre-training make sense?
von: Krones, Felix
Veröffentlicht: (2026)
von: Krones, Felix
Veröffentlicht: (2026)
CRISP: Hybrid Structured Sparsity for Class-aware Model Pruning
von: Aggarwal, Shivam, et al.
Veröffentlicht: (2023)
von: Aggarwal, Shivam, et al.
Veröffentlicht: (2023)
LOTUS: Improving Transformer Efficiency with Sparsity Pruning and Data Lottery Tickets
von: Upadhyay, Ojasw
Veröffentlicht: (2024)
von: Upadhyay, Ojasw
Veröffentlicht: (2024)
PLUM: Improving Inference Efficiency By Leveraging Repetition-Sparsity Trade-Off
von: Kuhar, Sachit, et al.
Veröffentlicht: (2023)
von: Kuhar, Sachit, et al.
Veröffentlicht: (2023)
Same accuracy, twice as fast: continuous training surpasses retraining from scratch
von: Verwimp, Eli, et al.
Veröffentlicht: (2025)
von: Verwimp, Eli, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Task agnostic continual learning with Pairwise layer architecture
von: Keskinen, Santtu
Veröffentlicht: (2024) -
Dynamic Sparse Training with Structured Sparsity
von: Lasby, Mike, et al.
Veröffentlicht: (2023) -
SINR: Sparsity Driven Compressed Implicit Neural Representations
von: Jayasundara, Dhananjaya, et al.
Veröffentlicht: (2025) -
LASERS: LAtent Space Encoding for Representations with Sparsity for Generative Modeling
von: Li, Xin, et al.
Veröffentlicht: (2024) -
Dual-Stage Invariant Continual Learning under Extreme Visual Sparsity
von: Zhang, Rangya, et al.
Veröffentlicht: (2026)