MeKi: Memory-based Expert Knowledge Injection for Efficient LLM Scaling
Fuente:
arXiv
Salvato in:
| Autori principali: | Ding, Ning, Liu, Fangcheng, Kim, Kyungrae, Hao, Linji, Lee, Kyeng-Hun, Ko, Hyeonmok, Tang, Yehui |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Efficient Compositional Multi-tasking for On-device Large Language Models
di: Bohdal, Ondrej, et al.
Pubblicazione: (2025)
di: Bohdal, Ondrej, et al.
Pubblicazione: (2025)
Data-driven Clustering and Merging of Adapters for On-device Large Language Models
di: Bohdal, Ondrej, et al.
Pubblicazione: (2026)
di: Bohdal, Ondrej, et al.
Pubblicazione: (2026)
HydraOpt: Navigating the Efficiency-Performance Trade-off of Adapter Merging
di: Ceritli, Taha, et al.
Pubblicazione: (2025)
di: Ceritli, Taha, et al.
Pubblicazione: (2025)
Hansel: Output Length Controlling Framework for Large Language Models
di: Song, Seoha, et al.
Pubblicazione: (2024)
di: Song, Seoha, et al.
Pubblicazione: (2024)
BRIDO: Bringing Democratic Order to Abstractive Summarization
di: Lee, Junhyun, et al.
Pubblicazione: (2025)
di: Lee, Junhyun, et al.
Pubblicazione: (2025)
On-device System of Compositional Multi-tasking in Large Language Models
di: Bohdal, Ondrej, et al.
Pubblicazione: (2025)
di: Bohdal, Ondrej, et al.
Pubblicazione: (2025)
Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning
di: Jie, Shibo, et al.
Pubblicazione: (2024)
di: Jie, Shibo, et al.
Pubblicazione: (2024)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
di: Long, Xiang, et al.
Pubblicazione: (2026)
di: Long, Xiang, et al.
Pubblicazione: (2026)
EAQuant: Enhancing Post-Training Quantization for MoE Models via Expert-Aware Optimization
di: Fu, Zhongqian, et al.
Pubblicazione: (2025)
di: Fu, Zhongqian, et al.
Pubblicazione: (2025)
Memory-Efficient Boundary Map for Large-Scale Occupancy Grid Mapping
di: Tang, Benxu, et al.
Pubblicazione: (2026)
di: Tang, Benxu, et al.
Pubblicazione: (2026)
FunHOI: Annotation-Free 3D Hand-Object Interaction Generation via Functional Text Guidanc
di: Tian, Yongqi, et al.
Pubblicazione: (2025)
di: Tian, Yongqi, et al.
Pubblicazione: (2025)
Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts
di: Rang, Miao, et al.
Pubblicazione: (2025)
di: Rang, Miao, et al.
Pubblicazione: (2025)
Kangaroo: Lossless Self-Speculative Decoding via Double Early Exiting
di: Liu, Fangcheng, et al.
Pubblicazione: (2024)
di: Liu, Fangcheng, et al.
Pubblicazione: (2024)
Personalized LLM Response Generation with Parameterized Memory Injection
di: Zhang, Kai, et al.
Pubblicazione: (2024)
di: Zhang, Kai, et al.
Pubblicazione: (2024)
Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
di: Bi, Zhenni, et al.
Pubblicazione: (2024)
di: Bi, Zhenni, et al.
Pubblicazione: (2024)
KiC: Keyword-inspired Cascade for Cost-Efficient Text Generation with LLMs
di: Kim, Woo-Chan, et al.
Pubblicazione: (2025)
di: Kim, Woo-Chan, et al.
Pubblicazione: (2025)
ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs
di: Ge, Hao, et al.
Pubblicazione: (2025)
di: Ge, Hao, et al.
Pubblicazione: (2025)
Mixture of Lookup Experts
di: Jie, Shibo, et al.
Pubblicazione: (2025)
di: Jie, Shibo, et al.
Pubblicazione: (2025)
MemoryFormer: Minimize Transformer Computation by Removing Fully-Connected Layers
di: Ding, Ning, et al.
Pubblicazione: (2024)
di: Ding, Ning, et al.
Pubblicazione: (2024)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
di: Liu, Xinyi, et al.
Pubblicazione: (2026)
di: Liu, Xinyi, et al.
Pubblicazione: (2026)
SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization
di: Lee, Wooin, et al.
Pubblicazione: (2026)
di: Lee, Wooin, et al.
Pubblicazione: (2026)
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
di: Jiang, Youhe, et al.
Pubblicazione: (2026)
di: Jiang, Youhe, et al.
Pubblicazione: (2026)
Differentiable Learning of Generalized Structured Matrices for Efficient Deep Neural Networks
di: Lee, Changwoo, et al.
Pubblicazione: (2023)
di: Lee, Changwoo, et al.
Pubblicazione: (2023)
Pangu Pro MoE: Mixture of Grouped Experts for Efficient Sparsity
di: Tang, Yehui, et al.
Pubblicazione: (2025)
di: Tang, Yehui, et al.
Pubblicazione: (2025)
DiffInject: Revisiting Debias via Synthetic Data Generation using Diffusion-based Style Injection
di: Ko, Donggeun, et al.
Pubblicazione: (2024)
di: Ko, Donggeun, et al.
Pubblicazione: (2024)
Hierarchical Knowledge Injection for Improving LLM-based Program Repair
di: Ehsani, Ramtin, et al.
Pubblicazione: (2025)
di: Ehsani, Ramtin, et al.
Pubblicazione: (2025)
IM-Chat: A Multi-agent LLM Framework Integrating Tool-Calling and Diffusion Modeling for Knowledge Transfer in Injection Molding Industry
di: Lee, Junhyeong, et al.
Pubblicazione: (2025)
di: Lee, Junhyeong, et al.
Pubblicazione: (2025)
Memory Injection Attacks on LLM Agents via Query-Only Interaction
di: Dong, Shen, et al.
Pubblicazione: (2025)
di: Dong, Shen, et al.
Pubblicazione: (2025)
SlimLLM: Accurate Structured Pruning for Large Language Models
di: Guo, Jialong, et al.
Pubblicazione: (2025)
di: Guo, Jialong, et al.
Pubblicazione: (2025)
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling
di: Ji, Xiaodong, et al.
Pubblicazione: (2025)
di: Ji, Xiaodong, et al.
Pubblicazione: (2025)
Thinking Short and Right Over Thinking Long: Serving LLM Reasoning Efficiently and Accurately
di: Wang, Yuhang, et al.
Pubblicazione: (2025)
di: Wang, Yuhang, et al.
Pubblicazione: (2025)
ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters
di: Hao, Zhiwei, et al.
Pubblicazione: (2025)
di: Hao, Zhiwei, et al.
Pubblicazione: (2025)
Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs
di: Han, Kai, et al.
Pubblicazione: (2024)
di: Han, Kai, et al.
Pubblicazione: (2024)
Memory-Efficient Structured Backpropagation for On-Device LLM Fine-Tuning
di: Park, Juneyoung, et al.
Pubblicazione: (2026)
di: Park, Juneyoung, et al.
Pubblicazione: (2026)
Low‐Temperature Photocrystallization of Atomic Layer Deposition‐Processed Tin Oxide for Highly Efficient and Flexible Perovskite Solar Cells
di: Dayeon Ko, et al.
Pubblicazione: (2025)
di: Dayeon Ko, et al.
Pubblicazione: (2025)
Memory-Efficient Acceleration of Block Low-Rank Foundation Models on Resource Constrained GPUs
di: Abillama, Pierre, et al.
Pubblicazione: (2025)
di: Abillama, Pierre, et al.
Pubblicazione: (2025)
GPT4Image: Large Pre-trained Models Help Vision Models Learn Better on Perception Task
di: Ding, Ning, et al.
Pubblicazione: (2023)
di: Ding, Ning, et al.
Pubblicazione: (2023)
Post-Training Quantization for Diffusion Transformer via Hierarchical Timestep Grouping
di: Ding, Ning, et al.
Pubblicazione: (2025)
di: Ding, Ning, et al.
Pubblicazione: (2025)
FlashRT: Towards Computationally and Memory Efficient Red-Teaming for Prompt Injection and Knowledge Corruption
di: Wang, Yanting, et al.
Pubblicazione: (2026)
di: Wang, Yanting, et al.
Pubblicazione: (2026)
Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge
di: Lu, Shuai, et al.
Pubblicazione: (2026)
di: Lu, Shuai, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Efficient Compositional Multi-tasking for On-device Large Language Models
di: Bohdal, Ondrej, et al.
Pubblicazione: (2025) -
Data-driven Clustering and Merging of Adapters for On-device Large Language Models
di: Bohdal, Ondrej, et al.
Pubblicazione: (2026) -
HydraOpt: Navigating the Efficiency-Performance Trade-off of Adapter Merging
di: Ceritli, Taha, et al.
Pubblicazione: (2025) -
Hansel: Output Length Controlling Framework for Large Language Models
di: Song, Seoha, et al.
Pubblicazione: (2024) -
BRIDO: Bringing Democratic Order to Abstractive Summarization
di: Lee, Junhyun, et al.
Pubblicazione: (2025)