Keep Your Friends Close: Leveraging Affinity Groups to Accelerate AI Inference Workflows
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Garrett, Thiago, Song, Weijia, Vitenberg, Roman, Birman, Ken |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Compass: A Decentralized Scheduler for Latency-Sensitive ML Workflows
von: Yang, Yuting, et al.
Veröffentlicht: (2024)
von: Yang, Yuting, et al.
Veröffentlicht: (2024)
On Replacing Cryptopuzzles with Useful Computation in Blockchain Proof-of-Work Protocols
von: Merlina, Andrea, et al.
Veröffentlicht: (2024)
von: Merlina, Andrea, et al.
Veröffentlicht: (2024)
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
A Discussion about Computational Challenges of Programmable Money in Blockchain-based CBDCs
von: da Conceição, Arlindo F., et al.
Veröffentlicht: (2024)
von: da Conceição, Arlindo F., et al.
Veröffentlicht: (2024)
Reconstruction-Based Adaptive Scheduling Using AI Inferences in Safety-Critical Systems
von: Alshaer, Samer, et al.
Veröffentlicht: (2025)
von: Alshaer, Samer, et al.
Veröffentlicht: (2025)
Accelerating LLM Inference with Precomputed Query Storage
von: Park, Jay H., et al.
Veröffentlicht: (2025)
von: Park, Jay H., et al.
Veröffentlicht: (2025)
The (R)evolution of Scientific Workflows in the Agentic AI Era: Towards Autonomous Science
von: Shin, Woong, et al.
Veröffentlicht: (2025)
von: Shin, Woong, et al.
Veröffentlicht: (2025)
Scalable AI-assisted Workflow Management for Detector Design Optimization Using Distributed Computing
von: Anderson, Derek, et al.
Veröffentlicht: (2026)
von: Anderson, Derek, et al.
Veröffentlicht: (2026)
Understand and Accelerate Memory Processing Pipeline for Large Language Model Inference
von: He, Zifan, et al.
Veröffentlicht: (2026)
von: He, Zifan, et al.
Veröffentlicht: (2026)
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)
Tangram: Accelerating Serverless LLM Loading through GPU Memory Reuse and Affinity
von: Zhu, Wenbin, et al.
Veröffentlicht: (2025)
von: Zhu, Wenbin, et al.
Veröffentlicht: (2025)
MSCCL++: Rethinking GPU Communication Abstractions for AI Inference
von: Hwang, Changho, et al.
Veröffentlicht: (2025)
von: Hwang, Changho, et al.
Veröffentlicht: (2025)
Decentralized AI: Permissionless LLM Inference on POKT Network
von: Olshansky, Daniel, et al.
Veröffentlicht: (2024)
von: Olshansky, Daniel, et al.
Veröffentlicht: (2024)
Leveraging Large Language Model for Intelligent Log Processing and Autonomous Debugging in Cloud AI Platforms
von: Ji, Cheng, et al.
Veröffentlicht: (2025)
von: Ji, Cheng, et al.
Veröffentlicht: (2025)
LegoDiffusion: Micro-Serving Text-to-Image Diffusion Workflows
von: Yang, Lingyun, et al.
Veröffentlicht: (2026)
von: Yang, Lingyun, et al.
Veröffentlicht: (2026)
Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2026)
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2026)
Accelerating a Triton Fused Kernel for W4A16 Quantized Inference with SplitK work decomposition
von: Hoque, Adnan, et al.
Veröffentlicht: (2024)
von: Hoque, Adnan, et al.
Veröffentlicht: (2024)
Accelerated Digital Twin Learning for Edge AI: A Comparison of FPGA and Mobile GPU
von: Xu, Bin, et al.
Veröffentlicht: (2025)
von: Xu, Bin, et al.
Veröffentlicht: (2025)
Cloud-Based AI Systems: Leveraging Large Language Models for Intelligent Fault Detection and Autonomous Self-Healing
von: Ji, Cheng, et al.
Veröffentlicht: (2025)
von: Ji, Cheng, et al.
Veröffentlicht: (2025)
AI Inference as Relocatable Electricity Demand: A Latency-Constrained Energy-Geography Framework
von: Luo, Xubin, et al.
Veröffentlicht: (2026)
von: Luo, Xubin, et al.
Veröffentlicht: (2026)
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
von: Song, Jingwei, et al.
Veröffentlicht: (2025)
von: Song, Jingwei, et al.
Veröffentlicht: (2025)
Reinforcement Learning-driven Data-intensive Workflow Scheduling for Volunteer Edge-Cloud
von: Mounesan, Motahare, et al.
Veröffentlicht: (2024)
von: Mounesan, Motahare, et al.
Veröffentlicht: (2024)
Electricity Cost Minimization for Multi-Workflow Allocation in Geo-Distributed Data Centers
von: Wang, Shuang, et al.
Veröffentlicht: (2025)
von: Wang, Shuang, et al.
Veröffentlicht: (2025)
WORKSWORLD: A Domain for Integrated Numeric Planning and Scheduling of Distributed Pipelined Workflows
von: Paul, Taylor, et al.
Veröffentlicht: (2026)
von: Paul, Taylor, et al.
Veröffentlicht: (2026)
Scalable Runtime Architecture for Data-driven, Hybrid HPC and ML Workflow Applications
von: Merzky, Andre, et al.
Veröffentlicht: (2025)
von: Merzky, Andre, et al.
Veröffentlicht: (2025)
DeServe: Towards Affordable Offline LLM Inference via Decentralization
von: Wu, Linyu, et al.
Veröffentlicht: (2025)
von: Wu, Linyu, et al.
Veröffentlicht: (2025)
Accelerating Latency-Critical Applications with AI-Powered Semi-Automatic Fine-Grained Parallelization on SMT Processors
von: Los, Denis, et al.
Veröffentlicht: (2025)
von: Los, Denis, et al.
Veröffentlicht: (2025)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025)
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025)
Characterizing Performance-Energy Trade-offs of Large Language Models in Multi-Request Workflows
von: Ifath, Md. Monzurul Amin, et al.
Veröffentlicht: (2026)
von: Ifath, Md. Monzurul Amin, et al.
Veröffentlicht: (2026)
KAITIAN: A Unified Communication Framework for Enabling Efficient Collaboration Across Heterogeneous Accelerators in Embodied AI Systems
von: Lin, Jieke, et al.
Veröffentlicht: (2025)
von: Lin, Jieke, et al.
Veröffentlicht: (2025)
Compass: Optimizing Compound AI Workflows for Dynamic Adaptation
von: Gravara, Milos, et al.
Veröffentlicht: (2026)
von: Gravara, Milos, et al.
Veröffentlicht: (2026)
Seesaw: High-throughput LLM Inference via Model Re-sharding
von: Su, Qidong, et al.
Veröffentlicht: (2025)
von: Su, Qidong, et al.
Veröffentlicht: (2025)
iOS as Acceleration
von: Chen, Alexander K.
Veröffentlicht: (2025)
von: Chen, Alexander K.
Veröffentlicht: (2025)
Accelerating MoE Model Inference with Expert Sharding
von: Balmau, Oana, et al.
Veröffentlicht: (2025)
von: Balmau, Oana, et al.
Veröffentlicht: (2025)
Mind the Gap: Revealing Inconsistencies Across Heterogeneous AI Accelerators
von: Wen, Elliott, et al.
Veröffentlicht: (2025)
von: Wen, Elliott, et al.
Veröffentlicht: (2025)
FIRST: Federated Inference Resource Scheduling Toolkit for Scientific AI Model Access
von: Tanikanti, Aditya, et al.
Veröffentlicht: (2025)
von: Tanikanti, Aditya, et al.
Veröffentlicht: (2025)
Multi-Agentic AI for Fairness-Aware and Accelerated Multi-modal Large Model Inference in Real-world Mobile Edge Networks
von: Li, Haiyuan, et al.
Veröffentlicht: (2026)
von: Li, Haiyuan, et al.
Veröffentlicht: (2026)
Characterizing Mobile SoC for Accelerating Heterogeneous LLM Inference
von: Chen, Le, et al.
Veröffentlicht: (2025)
von: Chen, Le, et al.
Veröffentlicht: (2025)
HadaCore: Tensor Core Accelerated Hadamard Transform Kernel
von: Agarwal, Krish, et al.
Veröffentlicht: (2024)
von: Agarwal, Krish, et al.
Veröffentlicht: (2024)
Beyond the Buzz: A Pragmatic Take on Inference Disaggregation
von: Mitra, Tiyasa, et al.
Veröffentlicht: (2025)
von: Mitra, Tiyasa, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Compass: A Decentralized Scheduler for Latency-Sensitive ML Workflows
von: Yang, Yuting, et al.
Veröffentlicht: (2024) -
On Replacing Cryptopuzzles with Useful Computation in Blockchain Proof-of-Work Protocols
von: Merlina, Andrea, et al.
Veröffentlicht: (2024) -
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
von: Yao, Jinghan, et al.
Veröffentlicht: (2024) -
A Discussion about Computational Challenges of Programmable Money in Blockchain-based CBDCs
von: da Conceição, Arlindo F., et al.
Veröffentlicht: (2024) -
Reconstruction-Based Adaptive Scheduling Using AI Inferences in Safety-Critical Systems
von: Alshaer, Samer, et al.
Veröffentlicht: (2025)