On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation
Fuente:
arXiv
Guardado en:
| Autores principales: | Diep, Nghiem T., Nguyen, Huy, Nguyen, Chau, Le, Minh, Nguyen, Duy M. H., Sonntag, Daniel, Niepert, Mathias, Ho, Nhat |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Hypernetwork-Driven Low-Rank Adaptation Across Attention Heads
por: Diep, Nghiem T., et al.
Publicado: (2025)
por: Diep, Nghiem T., et al.
Publicado: (2025)
Revisit Visual Prompt Tuning: The Expressiveness of Prompt Experts
por: Le, Minh, et al.
Publicado: (2025)
por: Le, Minh, et al.
Publicado: (2025)
Revisiting Prefix-tuning: Statistical Benefits of Reparameterization among Prompts
por: Le, Minh, et al.
Publicado: (2024)
por: Le, Minh, et al.
Publicado: (2024)
ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language Models
por: Nguyen, Duy M. H., et al.
Publicado: (2024)
por: Nguyen, Duy M. H., et al.
Publicado: (2024)
On Least Square Estimation in Softmax Gating Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2024)
por: Nguyen, Huy, et al.
Publicado: (2024)
Structure-Aware E(3)-Invariant Molecular Conformer Aggregation Networks
por: Nguyen, Duy M. H., et al.
Publicado: (2024)
por: Nguyen, Duy M. H., et al.
Publicado: (2024)
MGPATH: Vision-Language Model with Multi-Granular Prompt Learning for Few-Shot WSI Classification
por: Nguyen, Anh-Tien, et al.
Publicado: (2025)
por: Nguyen, Anh-Tien, et al.
Publicado: (2025)
RepLoRA: Reparameterizing Low-Rank Adaptation via the Perspective of Mixture of Experts
por: Truong, Tuan, et al.
Publicado: (2025)
por: Truong, Tuan, et al.
Publicado: (2025)
On Parameter Estimation in Deviated Gaussian Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2024)
por: Nguyen, Huy, et al.
Publicado: (2024)
Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2024)
por: Nguyen, Huy, et al.
Publicado: (2024)
Mixture of Experts Meets Prompt-Based Continual Learning
por: Le, Minh, et al.
Publicado: (2024)
por: Le, Minh, et al.
Publicado: (2024)
Quadratic Gating Mixture of Experts: Statistical Insights into Self-Attention
por: Akbarian, Pedram, et al.
Publicado: (2024)
por: Akbarian, Pedram, et al.
Publicado: (2024)
Accelerating Transformers with Spectrum-Preserving Token Merging
por: Tran, Hoai-Chau, et al.
Publicado: (2024)
por: Tran, Hoai-Chau, et al.
Publicado: (2024)
One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning
por: Le, Minh, et al.
Publicado: (2025)
por: Le, Minh, et al.
Publicado: (2025)
Convergence Rates for Softmax Gating Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2025)
por: Nguyen, Huy, et al.
Publicado: (2025)
A Statistical Theory of Gated Attention through the Lens of Hierarchical Mixture of Experts
por: Nguyen, Viet, et al.
Publicado: (2026)
por: Nguyen, Viet, et al.
Publicado: (2026)
DoRAN: Stabilizing Weight-Decomposed Low-Rank Adaptation via Noise Injection and Auxiliary Networks
por: Diep, Nghiem T., et al.
Publicado: (2025)
por: Diep, Nghiem T., et al.
Publicado: (2025)
A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2023)
por: Nguyen, Huy, et al.
Publicado: (2023)
Rethinking Multinomial Logistic Mixture of Experts with Sigmoid Gating Function
por: Pham, Tuan Minh, et al.
Publicado: (2026)
por: Pham, Tuan Minh, et al.
Publicado: (2026)
Towards Convergence Rates for Parameter Estimation in Gaussian-gated Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2023)
por: Nguyen, Huy, et al.
Publicado: (2023)
Dude: Dual Distribution-Aware Context Prompt Learning For Large Vision-Language Model
por: Nguyen, Duy M. H., et al.
Publicado: (2024)
por: Nguyen, Duy M. H., et al.
Publicado: (2024)
On Expert Estimation in Hierarchical Mixture of Experts: Beyond Softmax Gating Functions
por: Nguyen, Huy, et al.
Publicado: (2024)
por: Nguyen, Huy, et al.
Publicado: (2024)
Statistical Perspective of Top-K Sparse Softmax Gating Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2023)
por: Nguyen, Huy, et al.
Publicado: (2023)
S-Chain: Structured Visual Chain-of-Thought For Medicine
por: Le-Duc, Khai, et al.
Publicado: (2025)
por: Le-Duc, Khai, et al.
Publicado: (2025)
Fast Estimation of Wasserstein Distances via Regression on Sliced Wasserstein Distances
por: Nguyen, Khai, et al.
Publicado: (2025)
por: Nguyen, Khai, et al.
Publicado: (2025)
Is Temperature Sample Efficient for Softmax Gaussian Mixture of Experts?
por: Nguyen, Huy, et al.
Publicado: (2024)
por: Nguyen, Huy, et al.
Publicado: (2024)
On Bayesian Softmax-Gated Mixture-of-Experts Models
por: Bariletto, Nicola, et al.
Publicado: (2026)
por: Bariletto, Nicola, et al.
Publicado: (2026)
Attack On Prompt: Backdoor Attack in Prompt-Based Continual Learning
por: Nguyen, Trang, et al.
Publicado: (2024)
por: Nguyen, Trang, et al.
Publicado: (2024)
Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts
por: Yan, Fanqi, et al.
Publicado: (2024)
por: Yan, Fanqi, et al.
Publicado: (2024)
Improving Minimax Estimation Rates for Contaminated Mixture of Multinomial Logistic Experts via Expert Heterogeneity
por: Yan, Fanqi, et al.
Publicado: (2026)
por: Yan, Fanqi, et al.
Publicado: (2026)
Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling
por: Nguyen, Phuc Minh, et al.
Publicado: (2025)
por: Nguyen, Phuc Minh, et al.
Publicado: (2025)
On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts
por: Yan, Fanqi, et al.
Publicado: (2025)
por: Yan, Fanqi, et al.
Publicado: (2025)
Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective
por: Yan, Fanqi, et al.
Publicado: (2025)
por: Yan, Fanqi, et al.
Publicado: (2025)
Sliced Wasserstein Estimation with Control Variates
por: Nguyen, Khai, et al.
Publicado: (2023)
por: Nguyen, Khai, et al.
Publicado: (2023)
Adaptive Power Iteration Method for Differentially Private PCA
por: Nguyen, Ta Duy, et al.
Publicado: (2026)
por: Nguyen, Ta Duy, et al.
Publicado: (2026)
Lightspeed Geometric Dataset Distance via Sliced Optimal Transport
por: Nguyen, Khai, et al.
Publicado: (2025)
por: Nguyen, Khai, et al.
Publicado: (2025)
Towards Marginal Fairness Sliced Wasserstein Barycenter
por: Nguyen, Khai, et al.
Publicado: (2024)
por: Nguyen, Khai, et al.
Publicado: (2024)
BSO: Safety Alignment Is Density Ratio Matching
por: Nguyen, Tien-Phat, et al.
Publicado: (2026)
por: Nguyen, Tien-Phat, et al.
Publicado: (2026)
Statistical Advantages of Perturbing Cosine Router in Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2024)
por: Nguyen, Huy, et al.
Publicado: (2024)
On DeepSeekMoE: Statistical Benefits of Shared Experts and Normalized Sigmoid Gating
por: Nguyen, Huy, et al.
Publicado: (2025)
por: Nguyen, Huy, et al.
Publicado: (2025)
Ejemplares similares
-
Hypernetwork-Driven Low-Rank Adaptation Across Attention Heads
por: Diep, Nghiem T., et al.
Publicado: (2025) -
Revisit Visual Prompt Tuning: The Expressiveness of Prompt Experts
por: Le, Minh, et al.
Publicado: (2025) -
Revisiting Prefix-tuning: Statistical Benefits of Reparameterization among Prompts
por: Le, Minh, et al.
Publicado: (2024) -
ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language Models
por: Nguyen, Duy M. H., et al.
Publicado: (2024) -
On Least Square Estimation in Softmax Gating Mixture of Experts
por: Nguyen, Huy, et al.
Publicado: (2024)