Symbiosis: Multi-Adapter Inference and Fine-Tuning
Fuente:
arXiv
Guardado en:
| Autores principales: | Gupta, Saransh, Deshpande, Umesh, Janssen, Travis, Sundararaman, Swami |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
Fine-Tuning GPT-5 for GPU Kernel Generation
por: Tehrani, Ali, et al.
Publicado: (2026)
por: Tehrani, Ali, et al.
Publicado: (2026)
LoRAFusion: Efficient LoRA Fine-Tuning for LLMs
por: Zhu, Zhanda, et al.
Publicado: (2025)
por: Zhu, Zhanda, et al.
Publicado: (2025)
Enhancing Data Quality in Federated Fine-Tuning of Foundation Models
por: Zhao, Wanru, et al.
Publicado: (2024)
por: Zhao, Wanru, et al.
Publicado: (2024)
A Closer Look at Personalized Fine-Tuning in Heterogeneous Federated Learning
por: Chen, Minghui, et al.
Publicado: (2025)
por: Chen, Minghui, et al.
Publicado: (2025)
Efficient Federated Fine-Tuning of Large Language Models with Layer Dropout
por: Wang, Shilong, et al.
Publicado: (2025)
por: Wang, Shilong, et al.
Publicado: (2025)
Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
por: Li, Allison, et al.
Publicado: (2025)
por: Li, Allison, et al.
Publicado: (2025)
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
por: Sheng, Ying, et al.
Publicado: (2023)
por: Sheng, Ying, et al.
Publicado: (2023)
Learning Like Humans: Resource-Efficient Federated Fine-Tuning through Cognitive Developmental Stages
por: Wu, Yebo, et al.
Publicado: (2025)
por: Wu, Yebo, et al.
Publicado: (2025)
FLoRIST: Singular Value Thresholding for Efficient and Accurate Federated Fine-Tuning of Large Language Models
por: Ramesh, Hariharan, et al.
Publicado: (2025)
por: Ramesh, Hariharan, et al.
Publicado: (2025)
HSplitLoRA: A Heterogeneous Split Parameter-Efficient Fine-Tuning Framework for Large Language Models
por: Lin, Zheng, et al.
Publicado: (2025)
por: Lin, Zheng, et al.
Publicado: (2025)
Towards the Next Frontier of LLMs, Training on Private Data: A Cross-Domain Benchmark for Federated Fine-Tuning
por: Jimenez-Gutierrez, Daniel M., et al.
Publicado: (2026)
por: Jimenez-Gutierrez, Daniel M., et al.
Publicado: (2026)
FedRef: Bayesian Fine-Tuning using a Reference Model to Mitigate Catastrophic Forgetting for Heterogeneous Federated Learning
por: Yoon, Taehwan, et al.
Publicado: (2025)
por: Yoon, Taehwan, et al.
Publicado: (2025)
SFPrompt: Communication-Efficient Split Federated Fine-Tuning for Large Pre-Trained Models over Resource-Limited Devices
por: Cao, Linxiao, et al.
Publicado: (2024)
por: Cao, Linxiao, et al.
Publicado: (2024)
Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU
por: Levine, Reese, et al.
Publicado: (2026)
por: Levine, Reese, et al.
Publicado: (2026)
FSD-Inference: Fully Serverless Distributed Inference with Scalable Cloud Communication
por: Oakley, Joe, et al.
Publicado: (2024)
por: Oakley, Joe, et al.
Publicado: (2024)
FedPop: Federated Population-based Hyperparameter Tuning
por: Chen, Haokun, et al.
Publicado: (2023)
por: Chen, Haokun, et al.
Publicado: (2023)
Niyama : Breaking the Silos of LLM Inference Serving
por: Goel, Kanishk, et al.
Publicado: (2025)
por: Goel, Kanishk, et al.
Publicado: (2025)
On Evaluating Performance of LLM Inference Serving Systems
por: Agrawal, Amey, et al.
Publicado: (2025)
por: Agrawal, Amey, et al.
Publicado: (2025)
Stochastic Sparse Attention for Memory-Bound Inference
por: Lee, Kyle, et al.
Publicado: (2026)
por: Lee, Kyle, et al.
Publicado: (2026)
Leyline: KV Cache Directives for Agentic Inference
por: Ma, Bole, et al.
Publicado: (2026)
por: Ma, Bole, et al.
Publicado: (2026)
Lodestar: An Online-Learning LLM Inference Router
por: Lim, Gangmuk, et al.
Publicado: (2026)
por: Lim, Gangmuk, et al.
Publicado: (2026)
Context Parallelism for Scalable Million-Token Inference
por: Yang, Amy, et al.
Publicado: (2024)
por: Yang, Amy, et al.
Publicado: (2024)
DLoRA: Distributed Parameter-Efficient Fine-Tuning Solution for Large Language Model
por: Gao, Chao, et al.
Publicado: (2024)
por: Gao, Chao, et al.
Publicado: (2024)
Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
por: Zhang, Shulai, et al.
Publicado: (2025)
por: Zhang, Shulai, et al.
Publicado: (2025)
Accelerating MoE Model Inference with Expert Sharding
por: Balmau, Oana, et al.
Publicado: (2025)
por: Balmau, Oana, et al.
Publicado: (2025)
Frontier: Simulating the Next Generation of LLM Inference Systems
por: Feng, Yicheng, et al.
Publicado: (2025)
por: Feng, Yicheng, et al.
Publicado: (2025)
PIPO: Pipelined Offloading for Efficient Inference on Consumer Devices
por: Liu, Yangyijian, et al.
Publicado: (2025)
por: Liu, Yangyijian, et al.
Publicado: (2025)
Inference Offloading for Cost-Sensitive Binary Classification at the Edge
por: Moothedath, Vishnu Narayanan, et al.
Publicado: (2025)
por: Moothedath, Vishnu Narayanan, et al.
Publicado: (2025)
Frontier: Towards Comprehensive and Accurate LLM Inference Simulation
por: Feng, Yicheng, et al.
Publicado: (2026)
por: Feng, Yicheng, et al.
Publicado: (2026)
MatKV: Trading Compute for Flash Storage in LLM Inference
por: Shin, Kun-Woo, et al.
Publicado: (2025)
por: Shin, Kun-Woo, et al.
Publicado: (2025)
Characterizing Mobile SoC for Accelerating Heterogeneous LLM Inference
por: Chen, Le, et al.
Publicado: (2025)
por: Chen, Le, et al.
Publicado: (2025)
Adaptive Active Inference Agents for Heterogeneous and Lifelong Federated Learning
por: Danilenka, Anastasiya, et al.
Publicado: (2024)
por: Danilenka, Anastasiya, et al.
Publicado: (2024)
A Survey on Inference Optimization Techniques for Mixture of Experts Models
por: Liu, Jiacheng, et al.
Publicado: (2024)
por: Liu, Jiacheng, et al.
Publicado: (2024)
LLM-42: Enabling Determinism in LLM Inference with Verified Speculation
por: Gond, Raja, et al.
Publicado: (2026)
por: Gond, Raja, et al.
Publicado: (2026)
Efficient Fine-Grained GPU Performance Modeling for Distributed Deep Learning of LLM
por: Zhang, Biyao, et al.
Publicado: (2025)
por: Zhang, Biyao, et al.
Publicado: (2025)
HAFLQ: Heterogeneous Adaptive Federated LoRA Fine-tuned LLM with Quantization
por: Su, Yang, et al.
Publicado: (2024)
por: Su, Yang, et al.
Publicado: (2024)
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
por: Brüel-Gabrielsson, Rickard, et al.
Publicado: (2024)
por: Brüel-Gabrielsson, Rickard, et al.
Publicado: (2024)
Data Driven Optimization of GPU efficiency for Distributed LLM Adapter Serving
por: Agullo, Ferran, et al.
Publicado: (2026)
por: Agullo, Ferran, et al.
Publicado: (2026)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
por: Ye, Zihao, et al.
Publicado: (2025)
por: Ye, Zihao, et al.
Publicado: (2025)
Ejemplares similares
-
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
por: Jaiswal, Shashwat, et al.
Publicado: (2025) -
Fine-Tuning GPT-5 for GPU Kernel Generation
por: Tehrani, Ali, et al.
Publicado: (2026) -
LoRAFusion: Efficient LoRA Fine-Tuning for LLMs
por: Zhu, Zhanda, et al.
Publicado: (2025) -
Enhancing Data Quality in Federated Fine-Tuning of Foundation Models
por: Zhao, Wanru, et al.
Publicado: (2024) -
A Closer Look at Personalized Fine-Tuning in Heterogeneous Federated Learning
por: Chen, Minghui, et al.
Publicado: (2025)