A-IO: Adaptive Inference Orchestration for Memory-Bound NPUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Chen, Ding, Yan, Wang, Haotian, Liu, Chubo, Li, Keqin, Li, Kenli |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
OrchMLLM: Orchestrate Multimodal Data with Batch Post-Balancing to Accelerate Multimodal Large Language Model Training
von: Zheng, Yijie, et al.
Veröffentlicht: (2025)
von: Zheng, Yijie, et al.
Veröffentlicht: (2025)
Stochastic Sparse Attention for Memory-Bound Inference
von: Lee, Kyle, et al.
Veröffentlicht: (2026)
von: Lee, Kyle, et al.
Veröffentlicht: (2026)
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
von: Shen, Zixu, et al.
Veröffentlicht: (2025)
von: Shen, Zixu, et al.
Veröffentlicht: (2025)
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
von: Chen, Aodong, et al.
Veröffentlicht: (2023)
von: Chen, Aodong, et al.
Veröffentlicht: (2023)
Profiling-Driven Adaptive Distributed Transformer Inference on Embedded Edge Deployment
von: Qazi, Muhammad Azlan, et al.
Veröffentlicht: (2026)
von: Qazi, Muhammad Azlan, et al.
Veröffentlicht: (2026)
Verify Distributed Deep Learning Model Implementation Refinement with Iterative Relation Inference
von: Wang, Zhanghan, et al.
Veröffentlicht: (2025)
von: Wang, Zhanghan, et al.
Veröffentlicht: (2025)
Understand and Accelerate Memory Processing Pipeline for Large Language Model Inference
von: He, Zifan, et al.
Veröffentlicht: (2026)
von: He, Zifan, et al.
Veröffentlicht: (2026)
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
von: Liu, Wentao, et al.
Veröffentlicht: (2025)
von: Liu, Wentao, et al.
Veröffentlicht: (2025)
Data-Juicer 2.0: Cloud-Scale Adaptive Data Processing for and with Foundation Models
von: Chen, Daoyuan, et al.
Veröffentlicht: (2024)
von: Chen, Daoyuan, et al.
Veröffentlicht: (2024)
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
von: Suo, Jiashun, et al.
Veröffentlicht: (2025)
von: Suo, Jiashun, et al.
Veröffentlicht: (2025)
Failure-Resilient Distributed Inference with Model Compression over Heterogeneous Edge Devices
von: Wang, Li, et al.
Veröffentlicht: (2024)
von: Wang, Li, et al.
Veröffentlicht: (2024)
ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads
von: Zuo, Jingwei, et al.
Veröffentlicht: (2026)
von: Zuo, Jingwei, et al.
Veröffentlicht: (2026)
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
von: Bai, Fan, et al.
Veröffentlicht: (2026)
von: Bai, Fan, et al.
Veröffentlicht: (2026)
Reconstruction-Based Adaptive Scheduling Using AI Inferences in Safety-Critical Systems
von: Alshaer, Samer, et al.
Veröffentlicht: (2025)
von: Alshaer, Samer, et al.
Veröffentlicht: (2025)
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
von: Li, Xiangyu, et al.
Veröffentlicht: (2025)
von: Li, Xiangyu, et al.
Veröffentlicht: (2025)
Efficient Multi-Model Orchestration for Self-Hosted Large Language Models
von: Vangala, Bhanu Prakash, et al.
Veröffentlicht: (2025)
von: Vangala, Bhanu Prakash, et al.
Veröffentlicht: (2025)
Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management
von: Yang, Xinjun, et al.
Veröffentlicht: (2025)
von: Yang, Xinjun, et al.
Veröffentlicht: (2025)
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
von: Wu, Qi, et al.
Veröffentlicht: (2026)
von: Wu, Qi, et al.
Veröffentlicht: (2026)
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
Intelligent Autonomous Orchestration for Distributed Cloud Resources using Complex-Stability Analysis
von: Shyam, Gopal Krishna, et al.
Veröffentlicht: (2026)
von: Shyam, Gopal Krishna, et al.
Veröffentlicht: (2026)
xDiT: an Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
von: Fang, Jiarui, et al.
Veröffentlicht: (2024)
von: Fang, Jiarui, et al.
Veröffentlicht: (2024)
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
von: Cheng, Rongxin, et al.
Veröffentlicht: (2024)
von: Cheng, Rongxin, et al.
Veröffentlicht: (2024)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
von: Lyu, Hongtao, et al.
Veröffentlicht: (2025)
von: Lyu, Hongtao, et al.
Veröffentlicht: (2025)
Decentralized AI: Permissionless LLM Inference on POKT Network
von: Olshansky, Daniel, et al.
Veröffentlicht: (2024)
von: Olshansky, Daniel, et al.
Veröffentlicht: (2024)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
LLM Inference Serving: Survey of Recent Advances and Opportunities
von: Li, Baolin, et al.
Veröffentlicht: (2024)
von: Li, Baolin, et al.
Veröffentlicht: (2024)
Towards Carbon-Aware Container Orchestration: Predicting Workload Energy Consumption with Federated Learning
von: Saad, Zainab, et al.
Veröffentlicht: (2025)
von: Saad, Zainab, et al.
Veröffentlicht: (2025)
ScaleSim: Serving Large-Scale Multi-Agent Simulation with Invocation Distance-Based Memory Management
von: Pan, Zaifeng, et al.
Veröffentlicht: (2026)
von: Pan, Zaifeng, et al.
Veröffentlicht: (2026)
SparOA: Sparse and Operator-aware Hybrid Scheduling for Edge DNN Inference
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
DWDP: Distributed Weight Data Parallelism for High-Performance LLM Inference on NVL72
von: Li, Wanqian, et al.
Veröffentlicht: (2026)
von: Li, Wanqian, et al.
Veröffentlicht: (2026)
AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure
von: The AIBrix Team, et al.
Veröffentlicht: (2025)
von: The AIBrix Team, et al.
Veröffentlicht: (2025)
Domain-Adaptive Model Merging Across Disconnected Modes
von: Liu, Junming, et al.
Veröffentlicht: (2026)
von: Liu, Junming, et al.
Veröffentlicht: (2026)
KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
von: Zhang, Huawei, et al.
Veröffentlicht: (2025)
von: Zhang, Huawei, et al.
Veröffentlicht: (2025)
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)
PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep Learning
von: Wang, Yisu, et al.
Veröffentlicht: (2025)
von: Wang, Yisu, et al.
Veröffentlicht: (2025)
IslandRun: Privacy-Aware Multi-Objective Orchestration for Distributed AI Inference
von: Malepati, Bala Siva Sai Akhil
Veröffentlicht: (2025)
von: Malepati, Bala Siva Sai Akhil
Veröffentlicht: (2025)
HearthNet: Edge Multi-Agent Orchestration for Smart Homes
von: Zhan, Zhonghao, et al.
Veröffentlicht: (2026)
von: Zhan, Zhonghao, et al.
Veröffentlicht: (2026)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
von: Liu, Xing, et al.
Veröffentlicht: (2025)
von: Liu, Xing, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
von: Li, Zixuan, et al.
Veröffentlicht: (2024) -
OrchMLLM: Orchestrate Multimodal Data with Batch Post-Balancing to Accelerate Multimodal Large Language Model Training
von: Zheng, Yijie, et al.
Veröffentlicht: (2025) -
Stochastic Sparse Attention for Memory-Bound Inference
von: Lee, Kyle, et al.
Veröffentlicht: (2026) -
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
von: Shen, Zixu, et al.
Veröffentlicht: (2025) -
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
von: Chen, Aodong, et al.
Veröffentlicht: (2023)