Hypernetwork-Driven Low-Rank Adaptation Across Attention Heads
Fuente:
arXiv
Saved in:
| Main Authors: | Diep, Nghiem T., Le, Dung, Truong, Tuan, Dinh, Tan, Nguyen, Huy, Ho, Nhat |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DoRAN: Stabilizing Weight-Decomposed Low-Rank Adaptation via Noise Injection and Auxiliary Networks
by: Diep, Nghiem T., et al.
Published: (2025)
by: Diep, Nghiem T., et al.
Published: (2025)
RepLoRA: Reparameterizing Low-Rank Adaptation via the Perspective of Mixture of Experts
by: Truong, Tuan, et al.
Published: (2025)
by: Truong, Tuan, et al.
Published: (2025)
On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation
by: Diep, Nghiem T., et al.
Published: (2025)
by: Diep, Nghiem T., et al.
Published: (2025)
A Statistical Theory of Gated Attention through the Lens of Hierarchical Mixture of Experts
by: Nguyen, Viet, et al.
Published: (2026)
by: Nguyen, Viet, et al.
Published: (2026)
Improving Pareto Set Learning for Expensive Multi-objective Optimization via Stein Variational Hypernetworks
by: Nguyen, Minh-Duc, et al.
Published: (2024)
by: Nguyen, Minh-Duc, et al.
Published: (2024)
Improving Minimax Estimation Rates for Contaminated Mixture of Multinomial Logistic Experts via Expert Heterogeneity
by: Yan, Fanqi, et al.
Published: (2026)
by: Yan, Fanqi, et al.
Published: (2026)
Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts
by: Yan, Fanqi, et al.
Published: (2024)
by: Yan, Fanqi, et al.
Published: (2024)
A Framework for Controllable Multi-objective Learning with Annealed Stein Variational Hypernetworks
by: Nguyen, Minh-Duc, et al.
Published: (2025)
by: Nguyen, Minh-Duc, et al.
Published: (2025)
On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts
by: Yan, Fanqi, et al.
Published: (2025)
by: Yan, Fanqi, et al.
Published: (2025)
Improving Generalization with Flat Hilbert Bayesian Inference
by: Truong, Tuan, et al.
Published: (2024)
by: Truong, Tuan, et al.
Published: (2024)
Quadratic Gating Mixture of Experts: Statistical Insights into Self-Attention
by: Akbarian, Pedram, et al.
Published: (2024)
by: Akbarian, Pedram, et al.
Published: (2024)
On Parameter Estimation in Deviated Gaussian Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
Leveraging Hierarchical Taxonomies in Prompt-based Continual Learning
by: Tran, Quyen, et al.
Published: (2024)
by: Tran, Quyen, et al.
Published: (2024)
Stylized Text-to-Motion Generation via Hypernetwork-Driven Low-Rank Adaptation
by: Jeon, Junhyuk, et al.
Published: (2026)
by: Jeon, Junhyuk, et al.
Published: (2026)
Convergence Rates for Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2025)
by: Nguyen, Huy, et al.
Published: (2025)
On Least Square Estimation in Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
Is Temperature Sample Efficient for Softmax Gaussian Mixture of Experts?
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
Towards Efficient Pareto-optimal Utility-Fairness between Groups in Repeated Rankings
by: Mai, Phuong Dinh, et al.
Published: (2024)
by: Mai, Phuong Dinh, et al.
Published: (2024)
Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective
by: Yan, Fanqi, et al.
Published: (2025)
by: Yan, Fanqi, et al.
Published: (2025)
Rethinking Multinomial Logistic Mixture of Experts with Sigmoid Gating Function
by: Pham, Tuan Minh, et al.
Published: (2026)
by: Pham, Tuan Minh, et al.
Published: (2026)
Multi-Head Low-Rank Attention
by: Liu, Songtao, et al.
Published: (2026)
by: Liu, Songtao, et al.
Published: (2026)
Revisiting Prefix-tuning: Statistical Benefits of Reparameterization among Prompts
by: Le, Minh, et al.
Published: (2024)
by: Le, Minh, et al.
Published: (2024)
Deep Learning-Driven Friendly Jamming for Secure Multicarrier ISAC Under Channel Uncertainty
by: Tuan, Bui Minh, et al.
Published: (2026)
by: Tuan, Bui Minh, et al.
Published: (2026)
Towards Convergence Rates for Parameter Estimation in Gaussian-gated Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2023)
by: Nguyen, Huy, et al.
Published: (2023)
Promoting Ensemble Diversity with Interactive Bayesian Distributional Robustness for Fine-tuning Foundation Models
by: Pham, Ngoc-Quan, et al.
Published: (2025)
by: Pham, Ngoc-Quan, et al.
Published: (2025)
Channel Adaptation for EEG Foundation Models: A Systematic Benchmark Across Architectures, Tasks, and Training Regimes
by: Kokate, Kuntal, et al.
Published: (2026)
by: Kokate, Kuntal, et al.
Published: (2026)
iML: Executable, Problem-Grounded, and Broadly Exploratory Code-Driven AutoML
by: Le, Dat, et al.
Published: (2026)
by: Le, Dat, et al.
Published: (2026)
A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2023)
by: Nguyen, Huy, et al.
Published: (2023)
Low-Rank Adaptation of Neural Fields
by: Truong, Anh, et al.
Published: (2025)
by: Truong, Anh, et al.
Published: (2025)
Statistical Perspective of Top-K Sparse Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2023)
by: Nguyen, Huy, et al.
Published: (2023)
Hypernetworks for Perspectivist Adaptation
by: Ignatev, Daniil, et al.
Published: (2025)
by: Ignatev, Daniil, et al.
Published: (2025)
One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning
by: Le, Minh, et al.
Published: (2025)
by: Le, Minh, et al.
Published: (2025)
Mixture of Experts Meets Prompt-Based Continual Learning
by: Le, Minh, et al.
Published: (2024)
by: Le, Minh, et al.
Published: (2024)
Attention as a Hypernetwork
by: Schug, Simon, et al.
Published: (2024)
by: Schug, Simon, et al.
Published: (2024)
Revisit Visual Prompt Tuning: The Expressiveness of Prompt Experts
by: Le, Minh, et al.
Published: (2025)
by: Le, Minh, et al.
Published: (2025)
Data-Driven DRO and Economic Decision Theory: An Analytical Synthesis With Bayesian Nonparametric Advancements
by: Bariletto, Nicola, et al.
Published: (2024)
by: Bariletto, Nicola, et al.
Published: (2024)
Structure- and Stability-Preserving Learning of Port-Hamiltonian Systems
by: Nguyen, Binh, et al.
Published: (2026)
by: Nguyen, Binh, et al.
Published: (2026)
Lightspeed Geometric Dataset Distance via Sliced Optimal Transport
by: Nguyen, Khai, et al.
Published: (2025)
by: Nguyen, Khai, et al.
Published: (2025)
ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language Models
by: Nguyen, Duy M. H., et al.
Published: (2024)
by: Nguyen, Duy M. H., et al.
Published: (2024)
Similar Items
-
DoRAN: Stabilizing Weight-Decomposed Low-Rank Adaptation via Noise Injection and Auxiliary Networks
by: Diep, Nghiem T., et al.
Published: (2025) -
RepLoRA: Reparameterizing Low-Rank Adaptation via the Perspective of Mixture of Experts
by: Truong, Tuan, et al.
Published: (2025) -
On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation
by: Diep, Nghiem T., et al.
Published: (2025) -
A Statistical Theory of Gated Attention through the Lens of Hierarchical Mixture of Experts
by: Nguyen, Viet, et al.
Published: (2026) -
Improving Pareto Set Learning for Expensive Multi-objective Optimization via Stein Variational Hypernetworks
by: Nguyen, Minh-Duc, et al.
Published: (2024)