Symbiosis: Multi-Adapter Inference and Fine-Tuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gupta, Saransh, Deshpande, Umesh, Janssen, Travis, Sundararaman, Swami |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
Fine-Tuning GPT-5 for GPU Kernel Generation
von: Tehrani, Ali, et al.
Veröffentlicht: (2026)
von: Tehrani, Ali, et al.
Veröffentlicht: (2026)
LoRAFusion: Efficient LoRA Fine-Tuning for LLMs
von: Zhu, Zhanda, et al.
Veröffentlicht: (2025)
von: Zhu, Zhanda, et al.
Veröffentlicht: (2025)
Enhancing Data Quality in Federated Fine-Tuning of Foundation Models
von: Zhao, Wanru, et al.
Veröffentlicht: (2024)
von: Zhao, Wanru, et al.
Veröffentlicht: (2024)
A Closer Look at Personalized Fine-Tuning in Heterogeneous Federated Learning
von: Chen, Minghui, et al.
Veröffentlicht: (2025)
von: Chen, Minghui, et al.
Veröffentlicht: (2025)
Efficient Federated Fine-Tuning of Large Language Models with Layer Dropout
von: Wang, Shilong, et al.
Veröffentlicht: (2025)
von: Wang, Shilong, et al.
Veröffentlicht: (2025)
Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
von: Li, Allison, et al.
Veröffentlicht: (2025)
von: Li, Allison, et al.
Veröffentlicht: (2025)
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
von: Sheng, Ying, et al.
Veröffentlicht: (2023)
von: Sheng, Ying, et al.
Veröffentlicht: (2023)
Learning Like Humans: Resource-Efficient Federated Fine-Tuning through Cognitive Developmental Stages
von: Wu, Yebo, et al.
Veröffentlicht: (2025)
von: Wu, Yebo, et al.
Veröffentlicht: (2025)
FLoRIST: Singular Value Thresholding for Efficient and Accurate Federated Fine-Tuning of Large Language Models
von: Ramesh, Hariharan, et al.
Veröffentlicht: (2025)
von: Ramesh, Hariharan, et al.
Veröffentlicht: (2025)
HSplitLoRA: A Heterogeneous Split Parameter-Efficient Fine-Tuning Framework for Large Language Models
von: Lin, Zheng, et al.
Veröffentlicht: (2025)
von: Lin, Zheng, et al.
Veröffentlicht: (2025)
Towards the Next Frontier of LLMs, Training on Private Data: A Cross-Domain Benchmark for Federated Fine-Tuning
von: Jimenez-Gutierrez, Daniel M., et al.
Veröffentlicht: (2026)
von: Jimenez-Gutierrez, Daniel M., et al.
Veröffentlicht: (2026)
FedRef: Bayesian Fine-Tuning using a Reference Model to Mitigate Catastrophic Forgetting for Heterogeneous Federated Learning
von: Yoon, Taehwan, et al.
Veröffentlicht: (2025)
von: Yoon, Taehwan, et al.
Veröffentlicht: (2025)
SFPrompt: Communication-Efficient Split Federated Fine-Tuning for Large Pre-Trained Models over Resource-Limited Devices
von: Cao, Linxiao, et al.
Veröffentlicht: (2024)
von: Cao, Linxiao, et al.
Veröffentlicht: (2024)
Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU
von: Levine, Reese, et al.
Veröffentlicht: (2026)
von: Levine, Reese, et al.
Veröffentlicht: (2026)
FSD-Inference: Fully Serverless Distributed Inference with Scalable Cloud Communication
von: Oakley, Joe, et al.
Veröffentlicht: (2024)
von: Oakley, Joe, et al.
Veröffentlicht: (2024)
FedPop: Federated Population-based Hyperparameter Tuning
von: Chen, Haokun, et al.
Veröffentlicht: (2023)
von: Chen, Haokun, et al.
Veröffentlicht: (2023)
Niyama : Breaking the Silos of LLM Inference Serving
von: Goel, Kanishk, et al.
Veröffentlicht: (2025)
von: Goel, Kanishk, et al.
Veröffentlicht: (2025)
On Evaluating Performance of LLM Inference Serving Systems
von: Agrawal, Amey, et al.
Veröffentlicht: (2025)
von: Agrawal, Amey, et al.
Veröffentlicht: (2025)
Stochastic Sparse Attention for Memory-Bound Inference
von: Lee, Kyle, et al.
Veröffentlicht: (2026)
von: Lee, Kyle, et al.
Veröffentlicht: (2026)
Leyline: KV Cache Directives for Agentic Inference
von: Ma, Bole, et al.
Veröffentlicht: (2026)
von: Ma, Bole, et al.
Veröffentlicht: (2026)
Lodestar: An Online-Learning LLM Inference Router
von: Lim, Gangmuk, et al.
Veröffentlicht: (2026)
von: Lim, Gangmuk, et al.
Veröffentlicht: (2026)
Context Parallelism for Scalable Million-Token Inference
von: Yang, Amy, et al.
Veröffentlicht: (2024)
von: Yang, Amy, et al.
Veröffentlicht: (2024)
DLoRA: Distributed Parameter-Efficient Fine-Tuning Solution for Large Language Model
von: Gao, Chao, et al.
Veröffentlicht: (2024)
von: Gao, Chao, et al.
Veröffentlicht: (2024)
Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
von: Zhang, Shulai, et al.
Veröffentlicht: (2025)
von: Zhang, Shulai, et al.
Veröffentlicht: (2025)
Accelerating MoE Model Inference with Expert Sharding
von: Balmau, Oana, et al.
Veröffentlicht: (2025)
von: Balmau, Oana, et al.
Veröffentlicht: (2025)
Frontier: Simulating the Next Generation of LLM Inference Systems
von: Feng, Yicheng, et al.
Veröffentlicht: (2025)
von: Feng, Yicheng, et al.
Veröffentlicht: (2025)
PIPO: Pipelined Offloading for Efficient Inference on Consumer Devices
von: Liu, Yangyijian, et al.
Veröffentlicht: (2025)
von: Liu, Yangyijian, et al.
Veröffentlicht: (2025)
Inference Offloading for Cost-Sensitive Binary Classification at the Edge
von: Moothedath, Vishnu Narayanan, et al.
Veröffentlicht: (2025)
von: Moothedath, Vishnu Narayanan, et al.
Veröffentlicht: (2025)
Frontier: Towards Comprehensive and Accurate LLM Inference Simulation
von: Feng, Yicheng, et al.
Veröffentlicht: (2026)
von: Feng, Yicheng, et al.
Veröffentlicht: (2026)
MatKV: Trading Compute for Flash Storage in LLM Inference
von: Shin, Kun-Woo, et al.
Veröffentlicht: (2025)
von: Shin, Kun-Woo, et al.
Veröffentlicht: (2025)
Characterizing Mobile SoC for Accelerating Heterogeneous LLM Inference
von: Chen, Le, et al.
Veröffentlicht: (2025)
von: Chen, Le, et al.
Veröffentlicht: (2025)
Adaptive Active Inference Agents for Heterogeneous and Lifelong Federated Learning
von: Danilenka, Anastasiya, et al.
Veröffentlicht: (2024)
von: Danilenka, Anastasiya, et al.
Veröffentlicht: (2024)
A Survey on Inference Optimization Techniques for Mixture of Experts Models
von: Liu, Jiacheng, et al.
Veröffentlicht: (2024)
von: Liu, Jiacheng, et al.
Veröffentlicht: (2024)
LLM-42: Enabling Determinism in LLM Inference with Verified Speculation
von: Gond, Raja, et al.
Veröffentlicht: (2026)
von: Gond, Raja, et al.
Veröffentlicht: (2026)
Efficient Fine-Grained GPU Performance Modeling for Distributed Deep Learning of LLM
von: Zhang, Biyao, et al.
Veröffentlicht: (2025)
von: Zhang, Biyao, et al.
Veröffentlicht: (2025)
HAFLQ: Heterogeneous Adaptive Federated LoRA Fine-tuned LLM with Quantization
von: Su, Yang, et al.
Veröffentlicht: (2024)
von: Su, Yang, et al.
Veröffentlicht: (2024)
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
von: Brüel-Gabrielsson, Rickard, et al.
Veröffentlicht: (2024)
von: Brüel-Gabrielsson, Rickard, et al.
Veröffentlicht: (2024)
Data Driven Optimization of GPU efficiency for Distributed LLM Adapter Serving
von: Agullo, Ferran, et al.
Veröffentlicht: (2026)
von: Agullo, Ferran, et al.
Veröffentlicht: (2026)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
von: Ye, Zihao, et al.
Veröffentlicht: (2025)
von: Ye, Zihao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025) -
Fine-Tuning GPT-5 for GPU Kernel Generation
von: Tehrani, Ali, et al.
Veröffentlicht: (2026) -
LoRAFusion: Efficient LoRA Fine-Tuning for LLMs
von: Zhu, Zhanda, et al.
Veröffentlicht: (2025) -
Enhancing Data Quality in Federated Fine-Tuning of Foundation Models
von: Zhao, Wanru, et al.
Veröffentlicht: (2024) -
A Closer Look at Personalized Fine-Tuning in Heterogeneous Federated Learning
von: Chen, Minghui, et al.
Veröffentlicht: (2025)