Shears: Unstructured Sparsity with Neural Low-rank Adapter Search
Fuente:
arXiv
Salvato in:
| Autori principali: | Muñoz, J. Pablo, Yuan, Jinjie, Jain, Nilesh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Low-Rank Adapters Meet Neural Architecture Search for LLM Compression
di: Muñoz, J. Pablo, et al.
Pubblicazione: (2025)
di: Muñoz, J. Pablo, et al.
Pubblicazione: (2025)
SQFT: Low-cost Model Adaptation in Low-precision Sparse Foundation Models
di: Muñoz, Juan Pablo, et al.
Pubblicazione: (2024)
di: Muñoz, Juan Pablo, et al.
Pubblicazione: (2024)
Mamba-Shedder: Post-Transformer Compression for Efficient Selective Structured State Space Models
di: Muñoz, J. Pablo, et al.
Pubblicazione: (2025)
di: Muñoz, J. Pablo, et al.
Pubblicazione: (2025)
MultiPruner: Balanced Structure Removal in Foundation Models
di: Muñoz, J. Pablo, et al.
Pubblicazione: (2025)
di: Muñoz, J. Pablo, et al.
Pubblicazione: (2025)
SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning
di: Liao, Huanxuan, et al.
Pubblicazione: (2025)
di: Liao, Huanxuan, et al.
Pubblicazione: (2025)
OrchMoE: Efficient Multi-Adapter Learning with Task-Skill Synergy
di: Wang, Haowen, et al.
Pubblicazione: (2024)
di: Wang, Haowen, et al.
Pubblicazione: (2024)
Ensembles of Low-Rank Expert Adapters
di: Li, Yinghao, et al.
Pubblicazione: (2025)
di: Li, Yinghao, et al.
Pubblicazione: (2025)
KVCrush: Key value cache size-reduction using similarity in head-behaviour
di: Jha, Gopi Krishna, et al.
Pubblicazione: (2025)
di: Jha, Gopi Krishna, et al.
Pubblicazione: (2025)
LoRA-Pro: Are Low-Rank Adapters Properly Optimized?
di: Wang, Zhengbo, et al.
Pubblicazione: (2024)
di: Wang, Zhengbo, et al.
Pubblicazione: (2024)
Multiple Choice Learning of Low-Rank Adapters for Language Modeling
di: Letzelter, Victor, et al.
Pubblicazione: (2025)
di: Letzelter, Victor, et al.
Pubblicazione: (2025)
Training-Free Bayesianization for Low-Rank Adapters of Large Language Models
di: Shi, Haizhou, et al.
Pubblicazione: (2024)
di: Shi, Haizhou, et al.
Pubblicazione: (2024)
zFLoRA: Zero-Latency Fused Low-Rank Adapters
di: Gowda, Dhananjaya, et al.
Pubblicazione: (2025)
di: Gowda, Dhananjaya, et al.
Pubblicazione: (2025)
Post-Training Statistical Calibration for Higher Activation Sparsity
di: Chua, Vui Seng, et al.
Pubblicazione: (2024)
di: Chua, Vui Seng, et al.
Pubblicazione: (2024)
Low-rank finetuning for LLMs: A fairness perspective
di: Das, Saswat, et al.
Pubblicazione: (2024)
di: Das, Saswat, et al.
Pubblicazione: (2024)
Scalable LLM Reasoning Acceleration with Low-rank Distillation
di: Dong, Harry, et al.
Pubblicazione: (2025)
di: Dong, Harry, et al.
Pubblicazione: (2025)
Finer Parameter Steps for Low-Rank PEFT: A Controlled Study with CP Tensor Adapters
di: Wang, Xinjue, et al.
Pubblicazione: (2026)
di: Wang, Xinjue, et al.
Pubblicazione: (2026)
Solo Connection: A Parameter Efficient Fine-Tuning Technique for Transformers
di: Pathak, Harsh Nilesh, et al.
Pubblicazione: (2025)
di: Pathak, Harsh Nilesh, et al.
Pubblicazione: (2025)
Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration
di: Chen, Zhipeng, et al.
Pubblicazione: (2026)
di: Chen, Zhipeng, et al.
Pubblicazione: (2026)
Not All Adapters Matter: Selective Adapter Freezing for Memory-Efficient Fine-Tuning of Language Models
di: Son, Hyegang, et al.
Pubblicazione: (2024)
di: Son, Hyegang, et al.
Pubblicazione: (2024)
MoKA: Mixture of Kronecker Adapters
di: Sadeghi, Mohammadreza, et al.
Pubblicazione: (2025)
di: Sadeghi, Mohammadreza, et al.
Pubblicazione: (2025)
Random Initialization of Gated Sparse Adapters
di: Retault, Vi, et al.
Pubblicazione: (2025)
di: Retault, Vi, et al.
Pubblicazione: (2025)
The Role of Sparsity for Length Generalization in Transformers
di: Golowich, Noah, et al.
Pubblicazione: (2025)
di: Golowich, Noah, et al.
Pubblicazione: (2025)
Exploring Design Choices for Building Language-Specific LLMs
di: Tejaswi, Atula, et al.
Pubblicazione: (2024)
di: Tejaswi, Atula, et al.
Pubblicazione: (2024)
MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures
di: Ni, Jinjie, et al.
Pubblicazione: (2024)
di: Ni, Jinjie, et al.
Pubblicazione: (2024)
LegalLens: Leveraging LLMs for Legal Violation Identification in Unstructured Text
di: Bernsohn, Dor, et al.
Pubblicazione: (2024)
di: Bernsohn, Dor, et al.
Pubblicazione: (2024)
Dual-Personalizing Adapter for Federated Foundation Models
di: Yang, Yiyuan, et al.
Pubblicazione: (2024)
di: Yang, Yiyuan, et al.
Pubblicazione: (2024)
Neutral Residues: Revisiting Adapters for Model Extension
di: Talla, Franck Signe, et al.
Pubblicazione: (2024)
di: Talla, Franck Signe, et al.
Pubblicazione: (2024)
Learning Adapter Rank via Symmetry Breaking
di: Doyle, Cooper, et al.
Pubblicazione: (2025)
di: Doyle, Cooper, et al.
Pubblicazione: (2025)
Post-Training Sparse Attention with Double Sparsity
di: Yang, Shuo, et al.
Pubblicazione: (2024)
di: Yang, Shuo, et al.
Pubblicazione: (2024)
TokenButler: Token Importance is Predictable
di: Akhauri, Yash, et al.
Pubblicazione: (2025)
di: Akhauri, Yash, et al.
Pubblicazione: (2025)
Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
di: Minder, Julian, et al.
Pubblicazione: (2025)
di: Minder, Julian, et al.
Pubblicazione: (2025)
ABBA-Adapters: Efficient and Expressive Fine-Tuning of Foundation Models
di: Singhal, Raghav, et al.
Pubblicazione: (2025)
di: Singhal, Raghav, et al.
Pubblicazione: (2025)
Rethinking Adapter Placement: A Dominant Adaptation Module Perspective
di: Zhang, Suoxin, et al.
Pubblicazione: (2026)
di: Zhang, Suoxin, et al.
Pubblicazione: (2026)
Learn To be Efficient: Build Structured Sparsity in Large Language Models
di: Zheng, Haizhong, et al.
Pubblicazione: (2024)
di: Zheng, Haizhong, et al.
Pubblicazione: (2024)
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
di: Nakamura, Taishi, et al.
Pubblicazione: (2025)
di: Nakamura, Taishi, et al.
Pubblicazione: (2025)
SCALPEL: Selective Capability Ablation via Low-rank Parameter Editing for Large Language Model Interpretability Analysis
di: Fu, Zihao, et al.
Pubblicazione: (2026)
di: Fu, Zihao, et al.
Pubblicazione: (2026)
BBox-Adapter: Lightweight Adapting for Black-Box Large Language Models
di: Sun, Haotian, et al.
Pubblicazione: (2024)
di: Sun, Haotian, et al.
Pubblicazione: (2024)
Learning to Route for Dynamic Adapter Composition in Continual Learning with Language Models
di: Araujo, Vladimir, et al.
Pubblicazione: (2024)
di: Araujo, Vladimir, et al.
Pubblicazione: (2024)
Data-driven Clustering and Merging of Adapters for On-device Large Language Models
di: Bohdal, Ondrej, et al.
Pubblicazione: (2026)
di: Bohdal, Ondrej, et al.
Pubblicazione: (2026)
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
di: Akhauri, Yash, et al.
Pubblicazione: (2024)
di: Akhauri, Yash, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Low-Rank Adapters Meet Neural Architecture Search for LLM Compression
di: Muñoz, J. Pablo, et al.
Pubblicazione: (2025) -
SQFT: Low-cost Model Adaptation in Low-precision Sparse Foundation Models
di: Muñoz, Juan Pablo, et al.
Pubblicazione: (2024) -
Mamba-Shedder: Post-Transformer Compression for Efficient Selective Structured State Space Models
di: Muñoz, J. Pablo, et al.
Pubblicazione: (2025) -
MultiPruner: Balanced Structure Removal in Foundation Models
di: Muñoz, J. Pablo, et al.
Pubblicazione: (2025) -
SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning
di: Liao, Huanxuan, et al.
Pubblicazione: (2025)