Monkey Jump : MoE-Style PEFT for Efficient Multi-Task Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Prottasha, Nusrat Jahan, Kowsher, Md, Yu, Chun-Nam, Chen, Chen, Garibay, Ozlem |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LiME: Lightweight Mixture of Experts for Efficient Multimodal Multi-task Learning
von: Kowsher, Md, et al.
Veröffentlicht: (2026)
von: Kowsher, Md, et al.
Veröffentlicht: (2026)
Does Self-Attention Need Separate Weights in Transformers?
von: Kowsher, Md, et al.
Veröffentlicht: (2024)
von: Kowsher, Md, et al.
Veröffentlicht: (2024)
PEFT A2Z: Parameter-Efficient Fine-Tuning Survey for Large Language and Vision Models
von: Prottasha, Nusrat Jahan, et al.
Veröffentlicht: (2025)
von: Prottasha, Nusrat Jahan, et al.
Veröffentlicht: (2025)
Predicting Through Generation: Why Generation Is Better for Prediction
von: Kowsher, Md, et al.
Veröffentlicht: (2025)
von: Kowsher, Md, et al.
Veröffentlicht: (2025)
User Profile with Large Language Models: Construction, Updating, and Benchmarking
von: Prottasha, Nusrat Jahan, et al.
Veröffentlicht: (2025)
von: Prottasha, Nusrat Jahan, et al.
Veröffentlicht: (2025)
Propulsion: Steering LLM with Tiny Fine-Tuning
von: Kowsher, Md, et al.
Veröffentlicht: (2024)
von: Kowsher, Md, et al.
Veröffentlicht: (2024)
FlowNIB: An Information Bottleneck Analysis of Bidirectional vs. Unidirectional Language Models
von: Kowsher, Md, et al.
Veröffentlicht: (2025)
von: Kowsher, Md, et al.
Veröffentlicht: (2025)
Parameter-Efficient Fine-Tuning of Large Language Models using Semantic Knowledge Tuning
von: Prottasha, Nusrat Jahan, et al.
Veröffentlicht: (2024)
von: Prottasha, Nusrat Jahan, et al.
Veröffentlicht: (2024)
LLM-Mixer: Multiscale Mixing in LLMs for Time Series Forecasting
von: Kowsher, Md, et al.
Veröffentlicht: (2024)
von: Kowsher, Md, et al.
Veröffentlicht: (2024)
L-TUNING: Synchronized Label Tuning for Prompt and Prefix in LLMs
von: Kowsher, Md., et al.
Veröffentlicht: (2023)
von: Kowsher, Md., et al.
Veröffentlicht: (2023)
Token Trails: Navigating Contextual Depths in Conversational AI with ChatLLM
von: Kowsher, Md., et al.
Veröffentlicht: (2024)
von: Kowsher, Md., et al.
Veröffentlicht: (2024)
BnTTS: Few-Shot Speaker Adaptation in Low-Resource Setting
von: Basher, Mohammad Jahid Ibna, et al.
Veröffentlicht: (2025)
von: Basher, Mohammad Jahid Ibna, et al.
Veröffentlicht: (2025)
Octavius: Mitigating Task Interference in MLLMs via LoRA-MoE
von: Chen, Zeren, et al.
Veröffentlicht: (2023)
von: Chen, Zeren, et al.
Veröffentlicht: (2023)
RoCoFT: Efficient Finetuning of Large Language Models with Row-Column Updates
von: Kowsher, Md, et al.
Veröffentlicht: (2024)
von: Kowsher, Md, et al.
Veröffentlicht: (2024)
Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
Explainable Detection of Implicit Influential Patterns in Conversations via Data Augmentation
von: Abdidizaji, Sina, et al.
Veröffentlicht: (2025)
von: Abdidizaji, Sina, et al.
Veröffentlicht: (2025)
Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE
von: Chen, Junyi, et al.
Veröffentlicht: (2023)
von: Chen, Junyi, et al.
Veröffentlicht: (2023)
COME: Dual Structure-Semantic Learning with Collaborative MoE for Universal Lesion Detection Across Heterogeneous Ultrasound Datasets
von: Chen, Lingyu, et al.
Veröffentlicht: (2025)
von: Chen, Lingyu, et al.
Veröffentlicht: (2025)
SliceFine: The Universal Winning-Slice Hypothesis for Pretrained Networks
von: Kowsher, Md, et al.
Veröffentlicht: (2025)
von: Kowsher, Md, et al.
Veröffentlicht: (2025)
Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding
von: Liu, Xiangyue, et al.
Veröffentlicht: (2026)
von: Liu, Xiangyue, et al.
Veröffentlicht: (2026)
CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning
von: Liu, Yang, et al.
Veröffentlicht: (2026)
von: Liu, Yang, et al.
Veröffentlicht: (2026)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
von: Qian, Yulei, et al.
Veröffentlicht: (2024)
Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models
von: Jiang, Songtao, et al.
Veröffentlicht: (2024)
von: Jiang, Songtao, et al.
Veröffentlicht: (2024)
Pangu Pro MoE: Mixture of Grouped Experts for Efficient Sparsity
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
Expert Pyramid Tuning: Efficient Parameter Fine-Tuning for Expertise-Driven Task Allocation
von: Zhang, Jia-Chen, et al.
Veröffentlicht: (2026)
von: Zhang, Jia-Chen, et al.
Veröffentlicht: (2026)
Evaluating the Effectiveness of Cost-Efficient Large Language Models in Benchmark Biomedical Tasks
von: Jahan, Israt, et al.
Veröffentlicht: (2025)
von: Jahan, Israt, et al.
Veröffentlicht: (2025)
A Benchmark Study of Segmentation Models and Adaptation Strategies for Landslide Detection from Satellite Imagery
von: Kowsher, Md, et al.
Veröffentlicht: (2026)
von: Kowsher, Md, et al.
Veröffentlicht: (2026)
Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
MambaPEFT: Exploring Parameter-Efficient Fine-Tuning for Mamba
von: Yoshimura, Masakazu, et al.
Veröffentlicht: (2024)
von: Yoshimura, Masakazu, et al.
Veröffentlicht: (2024)
MoE-Sieve: Routing-Guided LoRA for Efficient MoE Fine-Tuning
von: Manzoni, Andrea
Veröffentlicht: (2026)
von: Manzoni, Andrea
Veröffentlicht: (2026)
3D-MoE: A Mixture-of-Experts Multi-modal LLM for 3D Vision and Pose Diffusion via Rectified Flow
von: Ma, Yueen, et al.
Veröffentlicht: (2025)
von: Ma, Yueen, et al.
Veröffentlicht: (2025)
Harder Tasks Need More Experts: Dynamic Routing in MoE Models
von: Huang, Quzhe, et al.
Veröffentlicht: (2024)
von: Huang, Quzhe, et al.
Veröffentlicht: (2024)
PMoL: Parameter Efficient MoE for Preference Mixing of LLM Alignment
von: Liu, Dongxu, et al.
Veröffentlicht: (2024)
von: Liu, Dongxu, et al.
Veröffentlicht: (2024)
MH-MoE: Multi-Head Mixture-of-Experts
von: Huang, Shaohan, et al.
Veröffentlicht: (2024)
von: Huang, Shaohan, et al.
Veröffentlicht: (2024)
Subject or Style: Adaptive and Training-Free Mixture of LoRAs
von: Zhang, Jia-Chen, et al.
Veröffentlicht: (2025)
von: Zhang, Jia-Chen, et al.
Veröffentlicht: (2025)
LLaDA-MoE: A Sparse MoE Diffusion Language Model
von: Zhu, Fengqi, et al.
Veröffentlicht: (2025)
von: Zhu, Fengqi, et al.
Veröffentlicht: (2025)
Courtroom-Style Multi-Agent Debate with Progressive RAG and Role-Switching for Controversial Claim Verification
von: Chowdhury, Masnun Nuha, et al.
Veröffentlicht: (2026)
von: Chowdhury, Masnun Nuha, et al.
Veröffentlicht: (2026)
Dense2MoE: Restructuring Diffusion Transformer to MoE for Efficient Text-to-Image Generation
von: Zheng, Youwei, et al.
Veröffentlicht: (2025)
von: Zheng, Youwei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LiME: Lightweight Mixture of Experts for Efficient Multimodal Multi-task Learning
von: Kowsher, Md, et al.
Veröffentlicht: (2026) -
Does Self-Attention Need Separate Weights in Transformers?
von: Kowsher, Md, et al.
Veröffentlicht: (2024) -
PEFT A2Z: Parameter-Efficient Fine-Tuning Survey for Large Language and Vision Models
von: Prottasha, Nusrat Jahan, et al.
Veröffentlicht: (2025) -
Predicting Through Generation: Why Generation Is Better for Prediction
von: Kowsher, Md, et al.
Veröffentlicht: (2025) -
User Profile with Large Language Models: Construction, Updating, and Benchmarking
von: Prottasha, Nusrat Jahan, et al.
Veröffentlicht: (2025)