Saved in:
| Main Authors: | Dun, Chen, Garcia, Mirian Hipolito, Zheng, Guoqing, Awadallah, Ahmed Hassan, Kyrillidis, Anastasios, Sim, Robert |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2310.02842 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings
by: Farhat, Yehya, et al.
Published: (2023)
by: Farhat, Yehya, et al.
Published: (2023)
Exploring How LLMs Capture and Represent Domain-Specific Knowledge
by: Garcia, Mirian Hipolito, et al.
Published: (2025)
by: Garcia, Mirian Hipolito, et al.
Published: (2025)
Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
by: Ding, Dujian, et al.
Published: (2024)
by: Ding, Dujian, et al.
Published: (2024)
Compressing LLMs with MoP: Mixture of Pruners
by: Yamamoto, Bruno Lopes, et al.
Published: (2026)
by: Yamamoto, Bruno Lopes, et al.
Published: (2026)
Guided by the Experts: Provable Feature Learning Dynamic of Soft-Routed Mixture-of-Experts
by: Liao, Fangshuo, et al.
Published: (2025)
by: Liao, Fangshuo, et al.
Published: (2025)
MoP-CLIP: A Mixture of Prompt-Tuned CLIP Models for Domain Incremental Learning
by: Nicolas, Julien, et al.
Published: (2023)
by: Nicolas, Julien, et al.
Published: (2023)
Assessing and Verifying Task Utility in LLM-Powered Applications
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Towards better Human-Agent Alignment: Assessing Task Utility in LLM-Powered Applications
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Creating Arabic LLM Prompts at Scale
by: El-Sheikh, Abdelrahman, et al.
Published: (2024)
by: El-Sheikh, Abdelrahman, et al.
Published: (2024)
GHOST: Unmasking Phantom States in Mamba2 via Grouped Hidden-state Output-aware Selection & Truncation
by: Menezes, Michael, et al.
Published: (2026)
by: Menezes, Michael, et al.
Published: (2026)
Provable Accelerated Convergence of Nesterov's Momentum for Deep ReLU Neural Networks
by: Liao, Fangshuo, et al.
Published: (2023)
by: Liao, Fangshuo, et al.
Published: (2023)
TAPO: Task-Referenced Adaptation for Prompt Optimization
by: Luo, Wenxin, et al.
Published: (2025)
by: Luo, Wenxin, et al.
Published: (2025)
Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents Through Help Feedback
by: Mehta, Nikhil, et al.
Published: (2023)
by: Mehta, Nikhil, et al.
Published: (2023)
Orca-Math: Unlocking the potential of SLMs in Grade School Math
by: Mitra, Arindam, et al.
Published: (2024)
by: Mitra, Arindam, et al.
Published: (2024)
MoR: Mixture of Ranks for Low-Rank Adaptation Tuning
by: Tang, Chuanyu, et al.
Published: (2024)
by: Tang, Chuanyu, et al.
Published: (2024)
MoIN: Mixture of Introvert Experts to Upcycle an LLM
by: Tejankar, Ajinkya, et al.
Published: (2024)
by: Tejankar, Ajinkya, et al.
Published: (2024)
MoPD: Mixture-of-Prompts Distillation for Vision-Language Models
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
Using non-convex optimization in quantum process tomography: Factored gradient descent is tough to beat
by: Quiroga, David A., et al.
Published: (2023)
by: Quiroga, David A., et al.
Published: (2023)
Better Schedules for Low Precision Training of Deep Neural Networks
by: Wolfe, Cameron R., et al.
Published: (2024)
by: Wolfe, Cameron R., et al.
Published: (2024)
PT-MoE: An Efficient Finetuning Framework for Integrating Mixture-of-Experts into Prompt Tuning
by: Li, Zongqian, et al.
Published: (2025)
by: Li, Zongqian, et al.
Published: (2025)
Prompt Attack Detection with LLM-as-a-Judge and Mixture-of-Models
by: Le, Hieu Xuan, et al.
Published: (2026)
by: Le, Hieu Xuan, et al.
Published: (2026)
LeMoLE: LLM-Enhanced Mixture of Linear Experts for Time Series Forecasting
by: Zhang, Lingzheng, et al.
Published: (2024)
by: Zhang, Lingzheng, et al.
Published: (2024)
OmniParser for Pure Vision Based GUI Agent
by: Lu, Yadong, et al.
Published: (2024)
by: Lu, Yadong, et al.
Published: (2024)
BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
by: Ding, Dujian, et al.
Published: (2025)
by: Ding, Dujian, et al.
Published: (2025)
The Geometry of Prompting: Unveiling Distinct Mechanisms of Task Adaptation in Language Models
by: Kirsanov, Artem, et al.
Published: (2025)
by: Kirsanov, Artem, et al.
Published: (2025)
Task Prompt Vectors: Effective Initialization through Multi-Task Soft-Prompt Transfer
by: Belanec, Robert, et al.
Published: (2024)
by: Belanec, Robert, et al.
Published: (2024)
MoA: Heterogeneous Mixture of Adapters for Parameter-Efficient Fine-Tuning of Large Language Models
by: Cao, Jie, et al.
Published: (2025)
by: Cao, Jie, et al.
Published: (2025)
Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use
by: Agarwal, Aradhye, et al.
Published: (2026)
by: Agarwal, Aradhye, et al.
Published: (2026)
MoRAgent: Parameter Efficient Agent Tuning with Mixture-of-Roles
by: Han, Jing, et al.
Published: (2025)
by: Han, Jing, et al.
Published: (2025)
Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for LLM Web Agents
by: Rosset, Corby, et al.
Published: (2024)
by: Rosset, Corby, et al.
Published: (2024)
NeuronMoE: Neuron-Guided Mixture-of-Experts for Efficient Multilingual LLM Extension
by: Li, Rongzhi, et al.
Published: (2026)
by: Li, Rongzhi, et al.
Published: (2026)
Dynamic Prompt Fusion for Multi-Task and Cross-Domain Adaptation in LLMs
by: Hu, Xin, et al.
Published: (2025)
by: Hu, Xin, et al.
Published: (2025)
MoECollab: Democratizing LLM Development Through Collaborative Mixture of Experts
by: Harshit
Published: (2025)
by: Harshit
Published: (2025)
One Model, Two Roles: Emergent Specialization in a Shared Recurrent Transformer
by: Shen, Jucheng, et al.
Published: (2026)
by: Shen, Jucheng, et al.
Published: (2026)
Quantum EigenGame for excited state calculation
by: Quiroga, David, et al.
Published: (2025)
by: Quiroga, David, et al.
Published: (2025)
Provable Model-Parallel Distributed Principal Component Analysis with Parallel Deflation
by: Liao, Fangshuo, et al.
Published: (2025)
by: Liao, Fangshuo, et al.
Published: (2025)
AdaPaD: Adaptive Parallel Deflation for PEFT with Self-Correcting Rank Discovery
by: Su, Barbara, et al.
Published: (2026)
by: Su, Barbara, et al.
Published: (2026)
SGD at the Edge of Stability: The Stochastic Sharpness Gap
by: Liao, Fangshuo, et al.
Published: (2026)
by: Liao, Fangshuo, et al.
Published: (2026)
Transfer-Prompting: Enhancing Cross-Task Adaptation in Large Language Models via Dual-Stage Prompts Optimization
by: Chang, Yupeng, et al.
Published: (2025)
by: Chang, Yupeng, et al.
Published: (2025)
Towards Specialized Generalists: A Multi-Task MoE-LoRA Framework for Domain-Specific LLM Adaptation
by: Yang, Yuxin, et al.
Published: (2026)
by: Yang, Yuxin, et al.
Published: (2026)
Similar Items
-
Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings
by: Farhat, Yehya, et al.
Published: (2023) -
Exploring How LLMs Capture and Represent Domain-Specific Knowledge
by: Garcia, Mirian Hipolito, et al.
Published: (2025) -
Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
by: Ding, Dujian, et al.
Published: (2024) -
Compressing LLMs with MoP: Mixture of Pruners
by: Yamamoto, Bruno Lopes, et al.
Published: (2026) -
Guided by the Experts: Provable Feature Learning Dynamic of Soft-Routed Mixture-of-Experts
by: Liao, Fangshuo, et al.
Published: (2025)