HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
Fuente:
arXiv
Salvato in:
| Autori principali: | Lin, Mao, Wang, Xi, Cox, Guilherme, Li, Dong, Jeon, Hyeran |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PASTA: A Modular Program Analysis Tool Framework for Accelerators
di: Lin, Mao, et al.
Pubblicazione: (2026)
di: Lin, Mao, et al.
Pubblicazione: (2026)
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
di: Li, Zhuojin, et al.
Pubblicazione: (2025)
di: Li, Zhuojin, et al.
Pubblicazione: (2025)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
di: Wahlgren, Jacob, et al.
Pubblicazione: (2025)
di: Wahlgren, Jacob, et al.
Pubblicazione: (2025)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
di: Vellaisamy, Prabhu, et al.
Pubblicazione: (2025)
di: Vellaisamy, Prabhu, et al.
Pubblicazione: (2025)
Efficient allocation of image recognition and LLM tasks on multi-GPU system
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)
Comparing CPU and GPU compute of PERMANOVA on MI300A
di: Sfiligoi, Igor
Pubblicazione: (2025)
di: Sfiligoi, Igor
Pubblicazione: (2025)
Multi-GPU Hybrid Particle-in-Cell Monte Carlo Simulations for Exascale Computing Systems
di: Williams, Jeremy J., et al.
Pubblicazione: (2026)
di: Williams, Jeremy J., et al.
Pubblicazione: (2026)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
di: Wang, Yuxin, et al.
Pubblicazione: (2023)
di: Wang, Yuxin, et al.
Pubblicazione: (2023)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
di: Davis, Joshua H., et al.
Pubblicazione: (2026)
di: Davis, Joshua H., et al.
Pubblicazione: (2026)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
di: Liu, Shifang, et al.
Pubblicazione: (2025)
di: Liu, Shifang, et al.
Pubblicazione: (2025)
Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures
di: Tharwani, Jay, et al.
Pubblicazione: (2025)
di: Tharwani, Jay, et al.
Pubblicazione: (2025)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024)
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
di: Zhang, Yaozheng, et al.
Pubblicazione: (2025)
di: Zhang, Yaozheng, et al.
Pubblicazione: (2025)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
di: Zhuang, Chen, et al.
Pubblicazione: (2024)
di: Zhuang, Chen, et al.
Pubblicazione: (2024)
Dissecting the software-based measurement of CPU energy consumption: a comparative analysis
di: Raffin, Guillaume, et al.
Pubblicazione: (2024)
di: Raffin, Guillaume, et al.
Pubblicazione: (2024)
Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC
di: Siavashi, Mohammad, et al.
Pubblicazione: (2026)
di: Siavashi, Mohammad, et al.
Pubblicazione: (2026)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
di: Zhang, Li, et al.
Pubblicazione: (2025)
di: Zhang, Li, et al.
Pubblicazione: (2025)
xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
di: Shi, Jiabo, et al.
Pubblicazione: (2025)
di: Shi, Jiabo, et al.
Pubblicazione: (2025)
Characterizing WebGPU Dispatch Overhead for LLM Inference Across Four GPU Vendors, Three Backends, and Three Browsers
di: Maczan, Jędrzej
Pubblicazione: (2026)
di: Maczan, Jędrzej
Pubblicazione: (2026)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
di: Karfakis, George, et al.
Pubblicazione: (2025)
di: Karfakis, George, et al.
Pubblicazione: (2025)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
di: Zhao, Yanbo, et al.
Pubblicazione: (2025)
di: Zhao, Yanbo, et al.
Pubblicazione: (2025)
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing
di: Xia, Yuning, et al.
Pubblicazione: (2026)
di: Xia, Yuning, et al.
Pubblicazione: (2026)
ADELIA: Automatic Differentiation for Efficient Laplace Inference Approximations
di: Boudaoud, Afif, et al.
Pubblicazione: (2026)
di: Boudaoud, Afif, et al.
Pubblicazione: (2026)
mLR: Scalable Laminography Reconstruction based on Memoization
di: Ma, Bin, et al.
Pubblicazione: (2025)
di: Ma, Bin, et al.
Pubblicazione: (2025)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
di: Dutt, Anurag, et al.
Pubblicazione: (2025)
di: Dutt, Anurag, et al.
Pubblicazione: (2025)
The Energy Cost of Execution-Idle in GPU Clusters
di: Lei, Yiran, et al.
Pubblicazione: (2026)
di: Lei, Yiran, et al.
Pubblicazione: (2026)
Scalable GPU Performance Variability Analysis framework
di: Lahiry, Ankur, et al.
Pubblicazione: (2025)
di: Lahiry, Ankur, et al.
Pubblicazione: (2025)
On the Partitioning of GPU Power among Multi-Instances
di: Vamja, Tirth, et al.
Pubblicazione: (2025)
di: Vamja, Tirth, et al.
Pubblicazione: (2025)
Unleashing the Power of Preemptive Priority-based Scheduling for Real-Time GPU Tasks
di: Wang, Yidi, et al.
Pubblicazione: (2024)
di: Wang, Yidi, et al.
Pubblicazione: (2024)
Disaggregated Design for GPU-Based Volumetric Data Structures
di: Meneghin, Massimiliano, et al.
Pubblicazione: (2025)
di: Meneghin, Massimiliano, et al.
Pubblicazione: (2025)
Taking GPU Programming Models to Task for Performance Portability
di: Davis, Joshua H., et al.
Pubblicazione: (2024)
di: Davis, Joshua H., et al.
Pubblicazione: (2024)
Profiling and optimization of multi-card GPU machine learning jobs
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
di: He, Jiaao, et al.
Pubblicazione: (2024)
di: He, Jiaao, et al.
Pubblicazione: (2024)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
di: Pilliat, Emmanuel
Pubblicazione: (2026)
di: Pilliat, Emmanuel
Pubblicazione: (2026)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
di: Islam, Tanzima Z., et al.
Pubblicazione: (2024)
di: Islam, Tanzima Z., et al.
Pubblicazione: (2024)
Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes
di: Mao, Ying, et al.
Pubblicazione: (2020)
di: Mao, Ying, et al.
Pubblicazione: (2020)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
di: Jain, Rutwik, et al.
Pubblicazione: (2026)
di: Jain, Rutwik, et al.
Pubblicazione: (2026)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
di: Mo, Zizhao, et al.
Pubblicazione: (2026)
di: Mo, Zizhao, et al.
Pubblicazione: (2026)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
di: Curless, Brian, et al.
Pubblicazione: (2025)
di: Curless, Brian, et al.
Pubblicazione: (2025)
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
di: Wang, Tuowei, et al.
Pubblicazione: (2024)
di: Wang, Tuowei, et al.
Pubblicazione: (2024)
Documenti analoghi
-
PASTA: A Modular Program Analysis Tool Framework for Accelerators
di: Lin, Mao, et al.
Pubblicazione: (2026) -
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
di: Li, Zhuojin, et al.
Pubblicazione: (2025) -
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
di: Wahlgren, Jacob, et al.
Pubblicazione: (2025) -
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
di: Vellaisamy, Prabhu, et al.
Pubblicazione: (2025) -
Efficient allocation of image recognition and LLM tasks on multi-GPU system
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)