MoKA: Mixture of Kronecker Adapters
Fuente:
arXiv
Saved in:
| Main Authors: | Sadeghi, Mohammadreza, Nejad, Mahsa Ghazvini, Asl, MirHamed Jafarzadeh, Gu, Yu, Yu, Yuanhao, Asgharian, Masoud, Nia, Vahid Partovi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HyWA: Hypernetwork Weight Adapting Personalized Voice Activity Detection
by: Nejad, Mahsa Ghazvini, et al.
Published: (2025)
by: Nejad, Mahsa Ghazvini, et al.
Published: (2025)
Tiny Noise-Robust Voice Activity Detector for Voice Assistants
by: Asl, Hamed Jafarzadeh, et al.
Published: (2025)
by: Asl, Hamed Jafarzadeh, et al.
Published: (2025)
Mitigating Outlier Activations in Low-Precision Fine-Tuning of Language Models
by: Ghaffari, Alireza, et al.
Published: (2023)
by: Ghaffari, Alireza, et al.
Published: (2023)
OAC: Output-adaptive Calibration for Accurate Post-training Quantization
by: Edalati, Ali, et al.
Published: (2024)
by: Edalati, Ali, et al.
Published: (2024)
Pause and Reflect: Conformal Aggregation for Chain-of-Thought Reasoning
by: Gu, Yu, et al.
Published: (2026)
by: Gu, Yu, et al.
Published: (2026)
AdpQ: A Zero-shot Calibration Free Adaptive Post Training Quantization Method for LLMs
by: Ghaffari, Alireza, et al.
Published: (2024)
by: Ghaffari, Alireza, et al.
Published: (2024)
$α$-Mutual Information: A Tunable Privacy Measure for Privacy Protection in Data Sharing
by: Asl, MirHamed Jafarzadeh, et al.
Published: (2023)
by: Asl, MirHamed Jafarzadeh, et al.
Published: (2023)
Rethinking Post-Training Quantization: Introducing a Statistical Pre-Calibration Approach
by: Ghaffari, Alireza, et al.
Published: (2025)
by: Ghaffari, Alireza, et al.
Published: (2025)
Understanding Neural Network Binarization with Forward and Backward Proximal Quantizers
by: Lu, Yiwei, et al.
Published: (2024)
by: Lu, Yiwei, et al.
Published: (2024)
PoTPTQ: A Two-step Power-of-Two Post-training for LLMs
by: Wang, Xinyu, et al.
Published: (2025)
by: Wang, Xinyu, et al.
Published: (2025)
Enterprise Resource Planning Using Multi-type Transformers in Ferro-Titanium Industry
by: Yazdanpourmoghadam, Samira, et al.
Published: (2026)
by: Yazdanpourmoghadam, Samira, et al.
Published: (2026)
Homa at SemEval-2025 Task 5: Aligning Librarian Records with OntoAligner for Subject Tagging
by: Tekanlou, Hadi Bayrami Asl, et al.
Published: (2025)
by: Tekanlou, Hadi Bayrami Asl, et al.
Published: (2025)
MoA: Heterogeneous Mixture of Adapters for Parameter-Efficient Fine-Tuning of Large Language Models
by: Cao, Jie, et al.
Published: (2025)
by: Cao, Jie, et al.
Published: (2025)
MoIN: Mixture of Introvert Experts to Upcycle an LLM
by: Tejankar, Ajinkya, et al.
Published: (2024)
by: Tejankar, Ajinkya, et al.
Published: (2024)
Introducing the Groundbreaking Design of the First n‐Input AND Gate Using an Optical‐MOSFET With a Revolutionary Structure of Pulse‐Shaped Gates
by: Hamed Emami‐Nejad, et al.
Published: (2026)
by: Hamed Emami‐Nejad, et al.
Published: (2026)
Structure-Learnable Adapter Fine-Tuning for Parameter-Efficient Large Language Models
by: Gong, Ming, et al.
Published: (2025)
by: Gong, Ming, et al.
Published: (2025)
Visualizing Spatial Point Clouds: A Task-Oriented Taxonomy
by: Partovi, Mahsa, et al.
Published: (2025)
by: Partovi, Mahsa, et al.
Published: (2025)
MING-MOE: Enhancing Medical Multi-Task Learning in Large Language Models with Sparse Mixture of Low-Rank Adapter Experts
by: Liao, Yusheng, et al.
Published: (2024)
by: Liao, Yusheng, et al.
Published: (2024)
OrchMoE: Efficient Multi-Adapter Learning with Task-Skill Synergy
by: Wang, Haowen, et al.
Published: (2024)
by: Wang, Haowen, et al.
Published: (2024)
MoM: Linear Sequence Modeling with Mixture-of-Memories
by: Du, Jusen, et al.
Published: (2025)
by: Du, Jusen, et al.
Published: (2025)
CuMA: Aligning LLMs with Sparse Cultural Values via Demographic-Aware Mixture of Adapters
by: Sun, Ao, et al.
Published: (2026)
by: Sun, Ao, et al.
Published: (2026)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
by: Gu, Naibin, et al.
Published: (2025)
by: Gu, Naibin, et al.
Published: (2025)
MoPD: Mixture-of-Prompts Distillation for Vision-Language Models
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
Sandwich attack: Multi-language Mixture Adaptive Attack on LLMs
by: Upadhayay, Bibek, et al.
Published: (2024)
by: Upadhayay, Bibek, et al.
Published: (2024)
Lightweight Zero-shot Text-to-Speech with Mixture of Adapters
by: Fujita, Kenichi, et al.
Published: (2024)
by: Fujita, Kenichi, et al.
Published: (2024)
Diving into Kronecker Adapters: Component Design Matters
by: Bai, Jiayu, et al.
Published: (2026)
by: Bai, Jiayu, et al.
Published: (2026)
KromHC: Manifold-Constrained Hyper-Connections with Kronecker-Product Residual Matrices
by: Zhou, Wuyang, et al.
Published: (2026)
by: Zhou, Wuyang, et al.
Published: (2026)
Mixture of insighTful Experts (MoTE): The Synergy of Thought Chains and Expert Mixtures in Self-Alignment
by: Liu, Zhili, et al.
Published: (2024)
by: Liu, Zhili, et al.
Published: (2024)
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models
by: Kang, Hao, et al.
Published: (2025)
by: Kang, Hao, et al.
Published: (2025)
$\texttt{MoE-RBench}$: Towards Building Reliable Language Models with Sparse Mixture-of-Experts
by: Chen, Guanjie, et al.
Published: (2024)
by: Chen, Guanjie, et al.
Published: (2024)
MoDEM: Mixture of Domain Expert Models
by: Simonds, Toby, et al.
Published: (2024)
by: Simonds, Toby, et al.
Published: (2024)
QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts
by: Li, Pingzhi, et al.
Published: (2024)
by: Li, Pingzhi, et al.
Published: (2024)
CLIP-Adapter: Better Vision-Language Models with Feature Adapters
by: Gao, Peng, et al.
Published: (2021)
by: Gao, Peng, et al.
Published: (2021)
Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation
by: Peng, Chunyi, et al.
Published: (2025)
by: Peng, Chunyi, et al.
Published: (2025)
Accelerated Gradient Methods for Sparse Statistical Learning with Nonconvex Penalties
by: Yang, Kai, et al.
Published: (2020)
by: Yang, Kai, et al.
Published: (2020)
MoLoRA: Composable Specialization via Per-Token Adapter Routing
by: Shah, Shrey, et al.
Published: (2026)
by: Shah, Shrey, et al.
Published: (2026)
Pangu Pro MoE: Mixture of Grouped Experts for Efficient Sparsity
by: Tang, Yehui, et al.
Published: (2025)
by: Tang, Yehui, et al.
Published: (2025)
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
by: Zhu, Tong, et al.
Published: (2024)
by: Zhu, Tong, et al.
Published: (2024)
DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts
by: Feng, Yuchen, et al.
Published: (2025)
by: Feng, Yuchen, et al.
Published: (2025)
Franken-Adapter: Cross-Lingual Adaptation of LLMs by Embedding Surgery
by: Jiang, Fan, et al.
Published: (2025)
by: Jiang, Fan, et al.
Published: (2025)
Similar Items
-
HyWA: Hypernetwork Weight Adapting Personalized Voice Activity Detection
by: Nejad, Mahsa Ghazvini, et al.
Published: (2025) -
Tiny Noise-Robust Voice Activity Detector for Voice Assistants
by: Asl, Hamed Jafarzadeh, et al.
Published: (2025) -
Mitigating Outlier Activations in Low-Precision Fine-Tuning of Language Models
by: Ghaffari, Alireza, et al.
Published: (2023) -
OAC: Output-adaptive Calibration for Accurate Post-training Quantization
by: Edalati, Ali, et al.
Published: (2024) -
Pause and Reflect: Conformal Aggregation for Chain-of-Thought Reasoning
by: Gu, Yu, et al.
Published: (2026)