Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kamahori, Keisuke, Tang, Tian, Gu, Yile, Zhu, Kan, Kasikci, Baris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ConsumerBench: Benchmarking Generative AI Applications on End-User Devices
von: Gu, Yile, et al.
Veröffentlicht: (2025)
von: Gu, Yile, et al.
Veröffentlicht: (2025)
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
Optimizing Task Scheduling in Heterogeneous Computing Environments: A Comparative Analysis of CPU, GPU, and ASIC Platforms Using E2C Simulator
von: Mohammadjafari, Ali, et al.
Veröffentlicht: (2024)
von: Mohammadjafari, Ali, et al.
Veröffentlicht: (2024)
Peformance Isolation for Inference Processes in Edge GPU Systems
von: Martín, Juan José, et al.
Veröffentlicht: (2026)
von: Martín, Juan José, et al.
Veröffentlicht: (2026)
VUDA: Breaking CUDA-Vulkan Isolation for Spatial Sharing of Compute and Graphics on the Same GPU
von: Xu, Bin, et al.
Veröffentlicht: (2026)
von: Xu, Bin, et al.
Veröffentlicht: (2026)
Dirigent: Lightweight Serverless Orchestration
von: Cvetković, Lazar, et al.
Veröffentlicht: (2024)
von: Cvetković, Lazar, et al.
Veröffentlicht: (2024)
PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated Speculation
von: Wei, Xingda, et al.
Veröffentlicht: (2024)
von: Wei, Xingda, et al.
Veröffentlicht: (2024)
Funky: Cloud-Native FPGA Virtualization and Orchestration
von: Koshiba, Atsushi, et al.
Veröffentlicht: (2025)
von: Koshiba, Atsushi, et al.
Veröffentlicht: (2025)
CPU-Limits kill Performance: Time to rethink Resource Control
von: Shetty, Chirag, et al.
Veröffentlicht: (2025)
von: Shetty, Chirag, et al.
Veröffentlicht: (2025)
Towards Efficient and Practical GPU Multitasking in the Era of LLM
von: Xing, Jiarong, et al.
Veröffentlicht: (2025)
von: Xing, Jiarong, et al.
Veröffentlicht: (2025)
Optimizing CPU Cache Utilization in Cloud VMs with Accurate Cache Abstraction
von: Tofigh, Mani, et al.
Veröffentlicht: (2025)
von: Tofigh, Mani, et al.
Veröffentlicht: (2025)
Performance Isolation and Semantic Determinism in Efficient GPU Spatial Sharing
von: Yang, Zhenyuan, et al.
Veröffentlicht: (2026)
von: Yang, Zhenyuan, et al.
Veröffentlicht: (2026)
NCCLbpf: Verified, Composable Policy Execution for GPU Collective Communication
von: Zheng, Yusheng
Veröffentlicht: (2026)
von: Zheng, Yusheng
Veröffentlicht: (2026)
Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC
von: Siavashi, Mohammad, et al.
Veröffentlicht: (2026)
von: Siavashi, Mohammad, et al.
Veröffentlicht: (2026)
GPUOS: A GPU Operating System Primitive for Transparent Operation Fusion
von: Yang, Yiwei, et al.
Veröffentlicht: (2026)
von: Yang, Yiwei, et al.
Veröffentlicht: (2026)
RAGDoll: Efficient Offloading-based Online RAG System on a Single GPU
von: Yu, Weiping, et al.
Veröffentlicht: (2025)
von: Yu, Weiping, et al.
Veröffentlicht: (2025)
BLITZSCALE: Fast and Live Large Model Autoscaling with O(1) Host Caching
von: Zhang, Dingyan, et al.
Veröffentlicht: (2024)
von: Zhang, Dingyan, et al.
Veröffentlicht: (2024)
FastMig: Leveraging FastFreeze to Establish Robust Service Liquidity in Cloud 2.0
von: Manatura, Sorawit, et al.
Veröffentlicht: (2024)
von: Manatura, Sorawit, et al.
Veröffentlicht: (2024)
GPUVM: GPU-driven Unified Virtual Memory
von: Nazaraliyev, Nurlan, et al.
Veröffentlicht: (2024)
von: Nazaraliyev, Nurlan, et al.
Veröffentlicht: (2024)
EdgeFlow: Fast Cold Starts for LLMs on Mobile Devices
von: Yan, Yongsheng, et al.
Veröffentlicht: (2026)
von: Yan, Yongsheng, et al.
Veröffentlicht: (2026)
Telepathic Datacenters: Fast RPCs using Shared CXL Memory
von: Mahar, Suyash, et al.
Veröffentlicht: (2024)
von: Mahar, Suyash, et al.
Veröffentlicht: (2024)
"Range as a Key" is the Key! Fast and Compact Cloud Block Store Index with RASK
von: Zhao, Haoru, et al.
Veröffentlicht: (2026)
von: Zhao, Haoru, et al.
Veröffentlicht: (2026)
NanoFlow: Towards Optimal Large Language Model Serving Throughput
von: Zhu, Kan, et al.
Veröffentlicht: (2024)
von: Zhu, Kan, et al.
Veröffentlicht: (2024)
Mercury: QoS-Aware Tiered Memory System
von: Lu, Jiaheng, et al.
Veröffentlicht: (2024)
von: Lu, Jiaheng, et al.
Veröffentlicht: (2024)
Formal Definitions and Performance Comparison of Consistency Models for Parallel File Systems
von: Wang, Chen, et al.
Veröffentlicht: (2024)
von: Wang, Chen, et al.
Veröffentlicht: (2024)
MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU
von: Yuan, Zhengqing, et al.
Veröffentlicht: (2026)
von: Yuan, Zhengqing, et al.
Veröffentlicht: (2026)
Vulcan: Instance-Optimal Systems Heuristics Through LLM-Driven Search
von: Dwivedula, Rohit, et al.
Veröffentlicht: (2025)
von: Dwivedula, Rohit, et al.
Veröffentlicht: (2025)
TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval
von: Lin, Chien-Yu, et al.
Veröffentlicht: (2025)
von: Lin, Chien-Yu, et al.
Veröffentlicht: (2025)
SwitchFS: Asynchronous Metadata Updates for Distributed Filesystems with In-Network Coordination
von: Xu, Jingwei, et al.
Veröffentlicht: (2024)
von: Xu, Jingwei, et al.
Veröffentlicht: (2024)
Samoyeds: Accelerating MoE Models with Structured Sparsity Leveraging Sparse Tensor Cores
von: Wu, Chenpeng, et al.
Veröffentlicht: (2025)
von: Wu, Chenpeng, et al.
Veröffentlicht: (2025)
Mitigating context switching in densely packed Linux clusters with Latency-Aware Group Scheduling
von: Isstaif, Al Amjad Tawfiq, et al.
Veröffentlicht: (2025)
von: Isstaif, Al Amjad Tawfiq, et al.
Veröffentlicht: (2025)
DPC: A Distributed Page Cache over CXL
von: Bergman, Shai, et al.
Veröffentlicht: (2026)
von: Bergman, Shai, et al.
Veröffentlicht: (2026)
Unlocking True Elasticity for the Cloud-Native Era with Dandelion
von: Kuchler, Tom, et al.
Veröffentlicht: (2025)
von: Kuchler, Tom, et al.
Veröffentlicht: (2025)
A Periodic Space of Distributed Computing: Vision & Framework
von: Salehi, Mohsen Amini, et al.
Veröffentlicht: (2026)
von: Salehi, Mohsen Amini, et al.
Veröffentlicht: (2026)
THEMIS: Time, Heterogeneity, and Energy Minded Scheduling for Fair Multi-Tenant Use in FPGAs
von: Karabulut, Emre, et al.
Veröffentlicht: (2024)
von: Karabulut, Emre, et al.
Veröffentlicht: (2024)
Mewz: Lightweight Execution Environment for WebAssembly with High Isolation and Portability using Unikernels
von: Ueda, Soichiro, et al.
Veröffentlicht: (2024)
von: Ueda, Soichiro, et al.
Veröffentlicht: (2024)
Taming Serverless Cold Starts Through OS Co-Design
von: Holmes, Ben, et al.
Veröffentlicht: (2025)
von: Holmes, Ben, et al.
Veröffentlicht: (2025)
Fix: externalizing network I/O in serverless computing
von: Deng, Yuhan, et al.
Veröffentlicht: (2025)
von: Deng, Yuhan, et al.
Veröffentlicht: (2025)
Equilibria: Fair Multi-Tenant CXL Memory Tiering At Scale
von: Zhao, Kaiyang, et al.
Veröffentlicht: (2026)
von: Zhao, Kaiyang, et al.
Veröffentlicht: (2026)
CvxCluster: Solving Large, Complex, Granular Resource Allocation Problems 100-1000x Faster
von: Nnorom Jr, Obi, et al.
Veröffentlicht: (2026)
von: Nnorom Jr, Obi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ConsumerBench: Benchmarking Generative AI Applications on End-User Devices
von: Gu, Yile, et al.
Veröffentlicht: (2025) -
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026) -
Optimizing Task Scheduling in Heterogeneous Computing Environments: A Comparative Analysis of CPU, GPU, and ASIC Platforms Using E2C Simulator
von: Mohammadjafari, Ali, et al.
Veröffentlicht: (2024) -
Peformance Isolation for Inference Processes in Edge GPU Systems
von: Martín, Juan José, et al.
Veröffentlicht: (2026) -
VUDA: Breaking CUDA-Vulkan Isolation for Spatial Sharing of Compute and Graphics on the Same GPU
von: Xu, Bin, et al.
Veröffentlicht: (2026)