Random Initialization of Gated Sparse Adapters
Fuente:
arXiv
Salvato in:
| Autori principali: | Retault, Vi, Berreby, Yohaï-Eliel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
JumpLoRA: Sparse Adapters for Continual Learning in Large Language Models
di: Dragomir, Alexandra, et al.
Pubblicazione: (2026)
di: Dragomir, Alexandra, et al.
Pubblicazione: (2026)
CuMA: Aligning LLMs with Sparse Cultural Values via Demographic-Aware Mixture of Adapters
di: Sun, Ao, et al.
Pubblicazione: (2026)
di: Sun, Ao, et al.
Pubblicazione: (2026)
Not All Adapters Matter: Selective Adapter Freezing for Memory-Efficient Fine-Tuning of Language Models
di: Son, Hyegang, et al.
Pubblicazione: (2024)
di: Son, Hyegang, et al.
Pubblicazione: (2024)
Ensembles of Low-Rank Expert Adapters
di: Li, Yinghao, et al.
Pubblicazione: (2025)
di: Li, Yinghao, et al.
Pubblicazione: (2025)
MoKA: Mixture of Kronecker Adapters
di: Sadeghi, Mohammadreza, et al.
Pubblicazione: (2025)
di: Sadeghi, Mohammadreza, et al.
Pubblicazione: (2025)
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
di: Riachi, Roland, et al.
Pubblicazione: (2025)
di: Riachi, Roland, et al.
Pubblicazione: (2025)
Learning Adapter Rank via Symmetry Breaking
di: Doyle, Cooper, et al.
Pubblicazione: (2025)
di: Doyle, Cooper, et al.
Pubblicazione: (2025)
Dual-Personalizing Adapter for Federated Foundation Models
di: Yang, Yiyuan, et al.
Pubblicazione: (2024)
di: Yang, Yiyuan, et al.
Pubblicazione: (2024)
Neutral Residues: Revisiting Adapters for Model Extension
di: Talla, Franck Signe, et al.
Pubblicazione: (2024)
di: Talla, Franck Signe, et al.
Pubblicazione: (2024)
Beyond Adapter Retrieval: Latent Geometry-Preserving Composition via Sparse Task Projection
di: Jin, Pengfei, et al.
Pubblicazione: (2024)
di: Jin, Pengfei, et al.
Pubblicazione: (2024)
ABBA-Adapters: Efficient and Expressive Fine-Tuning of Foundation Models
di: Singhal, Raghav, et al.
Pubblicazione: (2025)
di: Singhal, Raghav, et al.
Pubblicazione: (2025)
Multiple Choice Learning of Low-Rank Adapters for Language Modeling
di: Letzelter, Victor, et al.
Pubblicazione: (2025)
di: Letzelter, Victor, et al.
Pubblicazione: (2025)
LoRA-Pro: Are Low-Rank Adapters Properly Optimized?
di: Wang, Zhengbo, et al.
Pubblicazione: (2024)
di: Wang, Zhengbo, et al.
Pubblicazione: (2024)
Shears: Unstructured Sparsity with Neural Low-rank Adapter Search
di: Muñoz, J. Pablo, et al.
Pubblicazione: (2024)
di: Muñoz, J. Pablo, et al.
Pubblicazione: (2024)
Rethinking Adapter Placement: A Dominant Adaptation Module Perspective
di: Zhang, Suoxin, et al.
Pubblicazione: (2026)
di: Zhang, Suoxin, et al.
Pubblicazione: (2026)
Low-Rank Adapters Meet Neural Architecture Search for LLM Compression
di: Muñoz, J. Pablo, et al.
Pubblicazione: (2025)
di: Muñoz, J. Pablo, et al.
Pubblicazione: (2025)
zFLoRA: Zero-Latency Fused Low-Rank Adapters
di: Gowda, Dhananjaya, et al.
Pubblicazione: (2025)
di: Gowda, Dhananjaya, et al.
Pubblicazione: (2025)
Training-Free Bayesianization for Low-Rank Adapters of Large Language Models
di: Shi, Haizhou, et al.
Pubblicazione: (2024)
di: Shi, Haizhou, et al.
Pubblicazione: (2024)
Data-driven Clustering and Merging of Adapters for On-device Large Language Models
di: Bohdal, Ondrej, et al.
Pubblicazione: (2026)
di: Bohdal, Ondrej, et al.
Pubblicazione: (2026)
BBox-Adapter: Lightweight Adapting for Black-Box Large Language Models
di: Sun, Haotian, et al.
Pubblicazione: (2024)
di: Sun, Haotian, et al.
Pubblicazione: (2024)
Learning to Route for Dynamic Adapter Composition in Continual Learning with Language Models
di: Araujo, Vladimir, et al.
Pubblicazione: (2024)
di: Araujo, Vladimir, et al.
Pubblicazione: (2024)
Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
di: Krishna, Kundan, et al.
Pubblicazione: (2025)
di: Krishna, Kundan, et al.
Pubblicazione: (2025)
K-Merge: Online Continual Merging of Adapters for On-device Large Language Models
di: Shenaj, Donald, et al.
Pubblicazione: (2025)
di: Shenaj, Donald, et al.
Pubblicazione: (2025)
Generative Adapter: Contextualizing Language Models in Parameters with A Single Forward Pass
di: Chen, Tong, et al.
Pubblicazione: (2024)
di: Chen, Tong, et al.
Pubblicazione: (2024)
AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees
di: Fleshman, William, et al.
Pubblicazione: (2024)
di: Fleshman, William, et al.
Pubblicazione: (2024)
OrchMoE: Efficient Multi-Adapter Learning with Task-Skill Synergy
di: Wang, Haowen, et al.
Pubblicazione: (2024)
di: Wang, Haowen, et al.
Pubblicazione: (2024)
Task-Aware LoRA Adapter Composition via Similarity Retrieval in Vector Databases
di: Adsul, Riya, et al.
Pubblicazione: (2026)
di: Adsul, Riya, et al.
Pubblicazione: (2026)
Filter-then-Generate: Large Language Models with Structure-Text Adapter for Knowledge Graph Completion
di: Liu, Ben, et al.
Pubblicazione: (2024)
di: Liu, Ben, et al.
Pubblicazione: (2024)
MemAdapter: Fast Alignment across Agent Memory Paradigms via Generative Subgraph Retrieval
di: Zhang, Xin, et al.
Pubblicazione: (2026)
di: Zhang, Xin, et al.
Pubblicazione: (2026)
Finer Parameter Steps for Low-Rank PEFT: A Controlled Study with CP Tensor Adapters
di: Wang, Xinjue, et al.
Pubblicazione: (2026)
di: Wang, Xinjue, et al.
Pubblicazione: (2026)
Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs
di: Pepper, Keenan, et al.
Pubblicazione: (2026)
di: Pepper, Keenan, et al.
Pubblicazione: (2026)
MedualTime: A Dual-Adapter Language Model for Medical Time Series-Text Multimodal Learning
di: Ye, Jiexia, et al.
Pubblicazione: (2024)
di: Ye, Jiexia, et al.
Pubblicazione: (2024)
Instruction-Tuned, but Not More Verifiable Instruction-Following: A Cross-Task Diagnosis for LoRA Adapters
di: Zou, Junyi
Pubblicazione: (2026)
di: Zou, Junyi
Pubblicazione: (2026)
The Impact of Initialization on LoRA Finetuning Dynamics
di: Hayou, Soufiane, et al.
Pubblicazione: (2024)
di: Hayou, Soufiane, et al.
Pubblicazione: (2024)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
di: Yuan, Jingyang, et al.
Pubblicazione: (2025)
di: Yuan, Jingyang, et al.
Pubblicazione: (2025)
Merging Text Transformer Models from Different Initializations
di: Verma, Neha, et al.
Pubblicazione: (2024)
di: Verma, Neha, et al.
Pubblicazione: (2024)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
di: Chanin, David, et al.
Pubblicazione: (2025)
di: Chanin, David, et al.
Pubblicazione: (2025)
SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks
di: Christopoulou, Fenia, et al.
Pubblicazione: (2024)
di: Christopoulou, Fenia, et al.
Pubblicazione: (2024)
Forgetting Transformer: Softmax Attention with a Forget Gate
di: Lin, Zhixuan, et al.
Pubblicazione: (2025)
di: Lin, Zhixuan, et al.
Pubblicazione: (2025)
Grounded Token Initialization for New Vocabulary in LMs for Generative Recommendation
di: Chen, Daiwei, et al.
Pubblicazione: (2026)
di: Chen, Daiwei, et al.
Pubblicazione: (2026)
Documenti analoghi
-
JumpLoRA: Sparse Adapters for Continual Learning in Large Language Models
di: Dragomir, Alexandra, et al.
Pubblicazione: (2026) -
CuMA: Aligning LLMs with Sparse Cultural Values via Demographic-Aware Mixture of Adapters
di: Sun, Ao, et al.
Pubblicazione: (2026) -
Not All Adapters Matter: Selective Adapter Freezing for Memory-Efficient Fine-Tuning of Language Models
di: Son, Hyegang, et al.
Pubblicazione: (2024) -
Ensembles of Low-Rank Expert Adapters
di: Li, Yinghao, et al.
Pubblicazione: (2025) -
MoKA: Mixture of Kronecker Adapters
di: Sadeghi, Mohammadreza, et al.
Pubblicazione: (2025)