Practical offloading for fine-tuning LLM on commodity GPU via learned sparse projectors
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Siyuan, Wang, Zhuofeng, Guan, Zelong, Liu, Yudong, Gibbons, Phillip B. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
von: Chen, Siyuan, et al.
Veröffentlicht: (2025)
von: Chen, Siyuan, et al.
Veröffentlicht: (2025)
Carbon-aware decentralized dynamic task offloading in MIMO-MEC networks via multi-agent reinforcement learning
von: Zulfiqar, Mubshra, et al.
Veröffentlicht: (2026)
von: Zulfiqar, Mubshra, et al.
Veröffentlicht: (2026)
Fast LLM Post-training via Decoupled and Fastest-of-N Speculation
von: Cheng, Rongxin, et al.
Veröffentlicht: (2025)
von: Cheng, Rongxin, et al.
Veröffentlicht: (2025)
Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU
von: Spoczynski, Marcin, et al.
Veröffentlicht: (2026)
von: Spoczynski, Marcin, et al.
Veröffentlicht: (2026)
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
von: Liu, Yunzhao, et al.
Veröffentlicht: (2025)
von: Liu, Yunzhao, et al.
Veröffentlicht: (2025)
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
von: Argerich, Mauricio Fadel, et al.
Veröffentlicht: (2026)
von: Argerich, Mauricio Fadel, et al.
Veröffentlicht: (2026)
Task Graph offloading via Deep Reinforcement Learning in Mobile Edge Computing
von: Liu, Jiagang, et al.
Veröffentlicht: (2023)
von: Liu, Jiagang, et al.
Veröffentlicht: (2023)
An MLIR pipeline for offloading Fortran to FPGAs via OpenMP
von: Rodriguez-Canal, Gabriel, et al.
Veröffentlicht: (2025)
von: Rodriguez-Canal, Gabriel, et al.
Veröffentlicht: (2025)
Unified schemes for directive-based GPU offloading
von: Miki, Yohei, et al.
Veröffentlicht: (2024)
von: Miki, Yohei, et al.
Veröffentlicht: (2024)
Predictable LLM Serving on GPU Clusters
von: Darzi, Erfan, et al.
Veröffentlicht: (2025)
von: Darzi, Erfan, et al.
Veröffentlicht: (2025)
Towards Scalable GPU-Accelerated SNN Training via Temporal Fusion
von: Li, Yanchen, et al.
Veröffentlicht: (2024)
von: Li, Yanchen, et al.
Veröffentlicht: (2024)
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
von: Pan, Yudong, et al.
Veröffentlicht: (2026)
von: Pan, Yudong, et al.
Veröffentlicht: (2026)
Boosting LLM Serving through Spatial-Temporal GPU Resource Sharing
von: Lin, Zejia, et al.
Veröffentlicht: (2025)
von: Lin, Zejia, et al.
Veröffentlicht: (2025)
GPU-Virt-Bench: A Comprehensive Benchmarking Framework for Software-Based GPU Virtualization Systems
von: VG, Jithin, et al.
Veröffentlicht: (2025)
von: VG, Jithin, et al.
Veröffentlicht: (2025)
Towards Efficient and Practical GPU Multitasking in the Era of LLM
von: Xing, Jiarong, et al.
Veröffentlicht: (2025)
von: Xing, Jiarong, et al.
Veröffentlicht: (2025)
Cultivating Multidisciplinary AI Workforce Development on iTiger GPU Cluster: Practices and Challenges
von: Sharif, Mayira, et al.
Veröffentlicht: (2025)
von: Sharif, Mayira, et al.
Veröffentlicht: (2025)
PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers
von: Zhang, Hongbin, et al.
Veröffentlicht: (2026)
von: Zhang, Hongbin, et al.
Veröffentlicht: (2026)
A Practical GPU-Accelerated Implementation of Orthogonal Matching Pursuit
von: Lubonja, Ariel, et al.
Veröffentlicht: (2024)
von: Lubonja, Ariel, et al.
Veröffentlicht: (2024)
The 1/W Law: An Analytical Study of Context-Length Routing Topology and GPU Generation Gains for LLM Inference Energy Efficiency
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap
von: Qiang, Xinwei, et al.
Veröffentlicht: (2026)
von: Qiang, Xinwei, et al.
Veröffentlicht: (2026)
Cost-Efficient Multimodal LLM Inference via Cross-Tier GPU Heterogeneity
von: Yu, Donglin
Veröffentlicht: (2026)
von: Yu, Donglin
Veröffentlicht: (2026)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
von: Guo, Runsheng Benson, et al.
Veröffentlicht: (2025)
von: Guo, Runsheng Benson, et al.
Veröffentlicht: (2025)
UCCL-Zip: Lossless Compression Supercharged GPU Communication
von: Ma, Shuang, et al.
Veröffentlicht: (2026)
von: Ma, Shuang, et al.
Veröffentlicht: (2026)
MSCCL++: Rethinking GPU Communication Abstractions for AI Inference
von: Hwang, Changho, et al.
Veröffentlicht: (2025)
von: Hwang, Changho, et al.
Veröffentlicht: (2025)
Power- and Fragmentation-aware Online Scheduling for GPU Datacenters
von: Lettich, Francesco, et al.
Veröffentlicht: (2024)
von: Lettich, Francesco, et al.
Veröffentlicht: (2024)
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
von: Wu, Qi, et al.
Veröffentlicht: (2026)
von: Wu, Qi, et al.
Veröffentlicht: (2026)
TD-Orch: Scalable Load-Balancing for Distributed Systems with Applications to Graph Processing
von: Zhao, Yiwei, et al.
Veröffentlicht: (2025)
von: Zhao, Yiwei, et al.
Veröffentlicht: (2025)
LeMix: Unified Scheduling for LLM Training and Inference on Multi-GPU Systems
von: Li, Yufei, et al.
Veröffentlicht: (2025)
von: Li, Yufei, et al.
Veröffentlicht: (2025)
Tangram: Accelerating Serverless LLM Loading through GPU Memory Reuse and Affinity
von: Zhu, Wenbin, et al.
Veröffentlicht: (2025)
von: Zhu, Wenbin, et al.
Veröffentlicht: (2025)
Beyond the GPU: The Strategic Role of FPGAs in the Next Wave of AI
von: Jiménez, Arturo Urías
Veröffentlicht: (2025)
von: Jiménez, Arturo Urías
Veröffentlicht: (2025)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
von: Qianli, Liu, et al.
Veröffentlicht: (2025)
von: Qianli, Liu, et al.
Veröffentlicht: (2025)
DeServe: Towards Affordable Offline LLM Inference via Decentralization
von: Wu, Linyu, et al.
Veröffentlicht: (2025)
von: Wu, Linyu, et al.
Veröffentlicht: (2025)
KIS-S: A GPU-Aware Kubernetes Inference Simulator with RL-Based Auto-Scaling
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async
von: Liu, Yi, et al.
Veröffentlicht: (2025)
von: Liu, Yi, et al.
Veröffentlicht: (2025)
SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
von: Tschand, Arya, et al.
Veröffentlicht: (2025)
von: Tschand, Arya, et al.
Veröffentlicht: (2025)
Accelerating Large Language Model Training with Hybrid GPU-based Compression
von: Xu, Lang, et al.
Veröffentlicht: (2024)
von: Xu, Lang, et al.
Veröffentlicht: (2024)
An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU
von: Yang, Ruijia, et al.
Veröffentlicht: (2026)
von: Yang, Ruijia, et al.
Veröffentlicht: (2026)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2026)
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2026)
A sparsity-aware distributed-memory algorithm for sparse-sparse matrix multiplication
von: Hong, Yuxi, et al.
Veröffentlicht: (2024)
von: Hong, Yuxi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
von: Chen, Siyuan, et al.
Veröffentlicht: (2025) -
Carbon-aware decentralized dynamic task offloading in MIMO-MEC networks via multi-agent reinforcement learning
von: Zulfiqar, Mubshra, et al.
Veröffentlicht: (2026) -
Fast LLM Post-training via Decoupled and Fastest-of-N Speculation
von: Cheng, Rongxin, et al.
Veröffentlicht: (2025) -
Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU
von: Spoczynski, Marcin, et al.
Veröffentlicht: (2026) -
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
von: Liu, Yunzhao, et al.
Veröffentlicht: (2025)