RepLoRA: Reparameterizing Low-Rank Adaptation via the Perspective of Mixture of Experts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Truong, Tuan, Nguyen, Chau, Nguyen, Huy, Le, Minh, Le, Trung, Ho, Nhat |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Revisiting Prefix-tuning: Statistical Benefits of Reparameterization among Prompts
von: Le, Minh, et al.
Veröffentlicht: (2024)
von: Le, Minh, et al.
Veröffentlicht: (2024)
Hypernetwork-Driven Low-Rank Adaptation Across Attention Heads
von: Diep, Nghiem T., et al.
Veröffentlicht: (2025)
von: Diep, Nghiem T., et al.
Veröffentlicht: (2025)
Towards Convergence Rates for Parameter Estimation in Gaussian-gated Mixture of Experts
von: Nguyen, Huy, et al.
Veröffentlicht: (2023)
von: Nguyen, Huy, et al.
Veröffentlicht: (2023)
Mixture of Experts Meets Prompt-Based Continual Learning
von: Le, Minh, et al.
Veröffentlicht: (2024)
von: Le, Minh, et al.
Veröffentlicht: (2024)
A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts
von: Nguyen, Huy, et al.
Veröffentlicht: (2023)
von: Nguyen, Huy, et al.
Veröffentlicht: (2023)
Revisit Visual Prompt Tuning: The Expressiveness of Prompt Experts
von: Le, Minh, et al.
Veröffentlicht: (2025)
von: Le, Minh, et al.
Veröffentlicht: (2025)
On Parameter Estimation in Deviated Gaussian Mixture of Experts
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
Rethinking Multinomial Logistic Mixture of Experts with Sigmoid Gating Function
von: Pham, Tuan Minh, et al.
Veröffentlicht: (2026)
von: Pham, Tuan Minh, et al.
Veröffentlicht: (2026)
DoRAN: Stabilizing Weight-Decomposed Low-Rank Adaptation via Noise Injection and Auxiliary Networks
von: Diep, Nghiem T., et al.
Veröffentlicht: (2025)
von: Diep, Nghiem T., et al.
Veröffentlicht: (2025)
One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning
von: Le, Minh, et al.
Veröffentlicht: (2025)
von: Le, Minh, et al.
Veröffentlicht: (2025)
Improving Minimax Estimation Rates for Contaminated Mixture of Multinomial Logistic Experts via Expert Heterogeneity
von: Yan, Fanqi, et al.
Veröffentlicht: (2026)
von: Yan, Fanqi, et al.
Veröffentlicht: (2026)
Convergence Rates for Softmax Gating Mixture of Experts
von: Nguyen, Huy, et al.
Veröffentlicht: (2025)
von: Nguyen, Huy, et al.
Veröffentlicht: (2025)
Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts
von: Yan, Fanqi, et al.
Veröffentlicht: (2024)
von: Yan, Fanqi, et al.
Veröffentlicht: (2024)
Statistical Perspective of Top-K Sparse Softmax Gating Mixture of Experts
von: Nguyen, Huy, et al.
Veröffentlicht: (2023)
von: Nguyen, Huy, et al.
Veröffentlicht: (2023)
On Least Square Estimation in Softmax Gating Mixture of Experts
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
Is Temperature Sample Efficient for Softmax Gaussian Mixture of Experts?
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
A Statistical Theory of Gated Attention through the Lens of Hierarchical Mixture of Experts
von: Nguyen, Viet, et al.
Veröffentlicht: (2026)
von: Nguyen, Viet, et al.
Veröffentlicht: (2026)
Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts
von: Yan, Fanqi, et al.
Veröffentlicht: (2025)
von: Yan, Fanqi, et al.
Veröffentlicht: (2025)
Quadratic Gating Mixture of Experts: Statistical Insights into Self-Attention
von: Akbarian, Pedram, et al.
Veröffentlicht: (2024)
von: Akbarian, Pedram, et al.
Veröffentlicht: (2024)
On Expert Estimation in Hierarchical Mixture of Experts: Beyond Softmax Gating Functions
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
Statistical Advantages of Perturbing Cosine Router in Mixture of Experts
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
Model Selection for Gaussian-gated Gaussian Mixture of Experts Using Dendrograms of Mixing Measures
von: Thai, Tuan, et al.
Veröffentlicht: (2025)
von: Thai, Tuan, et al.
Veröffentlicht: (2025)
Leveraging Hierarchical Taxonomies in Prompt-based Continual Learning
von: Tran, Quyen, et al.
Veröffentlicht: (2024)
von: Tran, Quyen, et al.
Veröffentlicht: (2024)
CompeteSMoE -- Effective Training of Sparse Mixture of Experts via Competition
von: Pham, Quang, et al.
Veröffentlicht: (2024)
von: Pham, Quang, et al.
Veröffentlicht: (2024)
FuseMoE: Mixture-of-Experts Transformers for Fleximodal Fusion
von: Han, Xing, et al.
Veröffentlicht: (2024)
von: Han, Xing, et al.
Veröffentlicht: (2024)
On Bayesian Softmax-Gated Mixture-of-Experts Models
von: Bariletto, Nicola, et al.
Veröffentlicht: (2026)
von: Bariletto, Nicola, et al.
Veröffentlicht: (2026)
Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective
von: Yan, Fanqi, et al.
Veröffentlicht: (2025)
von: Yan, Fanqi, et al.
Veröffentlicht: (2025)
Dendrograms of Mixing Measures for Softmax-Gated Gaussian Mixture of Experts: Consistency without Model Sweeps
von: Hai, Do Tien, et al.
Veröffentlicht: (2025)
von: Hai, Do Tien, et al.
Veröffentlicht: (2025)
On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation
von: Diep, Nghiem T., et al.
Veröffentlicht: (2025)
von: Diep, Nghiem T., et al.
Veröffentlicht: (2025)
MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
von: Huynh, Tuan-Luc, et al.
Veröffentlicht: (2025)
von: Huynh, Tuan-Luc, et al.
Veröffentlicht: (2025)
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
von: Nguyen-Nhat, Minh-Khoi, et al.
Veröffentlicht: (2025)
von: Nguyen-Nhat, Minh-Khoi, et al.
Veröffentlicht: (2025)
HELLoRA: Hot Experts Layer-Level Low-Rank Adaptation for Mixture-of-Experts Models
von: Wei, Jia, et al.
Veröffentlicht: (2026)
von: Wei, Jia, et al.
Veröffentlicht: (2026)
Queryable LoRA: Instruction-Regularized Routing Over Shared Low-Rank Update Atoms
von: Vaidya, Omatharv Bharat, et al.
Veröffentlicht: (2026)
von: Vaidya, Omatharv Bharat, et al.
Veröffentlicht: (2026)
Spectral Flattening Is All Muon Needs: How Orthogonalization Controls Learning Rate and Convergence
von: Nguyen, Tien-Phat, et al.
Veröffentlicht: (2026)
von: Nguyen, Tien-Phat, et al.
Veröffentlicht: (2026)
Improving Generalization with Flat Hilbert Bayesian Inference
von: Truong, Tuan, et al.
Veröffentlicht: (2024)
von: Truong, Tuan, et al.
Veröffentlicht: (2024)
BSO: Safety Alignment Is Density Ratio Matching
von: Nguyen, Tien-Phat, et al.
Veröffentlicht: (2026)
von: Nguyen, Tien-Phat, et al.
Veröffentlicht: (2026)
Fast Model Selection and Stable Optimization for Softmax-Gated Multinomial-Logistic Mixture of Experts Models
von: Tran, TrungKhang, et al.
Veröffentlicht: (2026)
von: Tran, TrungKhang, et al.
Veröffentlicht: (2026)
Mixture of LoRA Experts
von: Wu, Xun, et al.
Veröffentlicht: (2024)
von: Wu, Xun, et al.
Veröffentlicht: (2024)
Beyond Losses Reweighting: Empowering Multi-Task Learning via the Generalization Perspective
von: Phan, Hoang, et al.
Veröffentlicht: (2022)
von: Phan, Hoang, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
Revisiting Prefix-tuning: Statistical Benefits of Reparameterization among Prompts
von: Le, Minh, et al.
Veröffentlicht: (2024) -
Hypernetwork-Driven Low-Rank Adaptation Across Attention Heads
von: Diep, Nghiem T., et al.
Veröffentlicht: (2025) -
Towards Convergence Rates for Parameter Estimation in Gaussian-gated Mixture of Experts
von: Nguyen, Huy, et al.
Veröffentlicht: (2023) -
Mixture of Experts Meets Prompt-Based Continual Learning
von: Le, Minh, et al.
Veröffentlicht: (2024) -
A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts
von: Nguyen, Huy, et al.
Veröffentlicht: (2023)