Dynamic Rebatching for Efficient Early-Exit Inference with DREX
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Xuting, Alexander, Daniel, Kakarla, Siva Kesava Reddy, Arzani, Behnaz, Liu, Vincent |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Performance Analyzer for a Public Cloud's ML-Augmented VM Allocator
di: Bostandoost, Roozbeh, et al.
Pubblicazione: (2025)
di: Bostandoost, Roozbeh, et al.
Pubblicazione: (2025)
Towards Safer Heuristics With XPlain
di: Karimi, Pantea, et al.
Pubblicazione: (2024)
di: Karimi, Pantea, et al.
Pubblicazione: (2024)
Federated Learning for Collaborative Inference Systems: The Case of Early Exit Networks
di: Kaplan, Caelin, et al.
Pubblicazione: (2024)
di: Kaplan, Caelin, et al.
Pubblicazione: (2024)
Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving
di: Dai, Yinwei, et al.
Pubblicazione: (2023)
di: Dai, Yinwei, et al.
Pubblicazione: (2023)
Distributed Inference on Mobile Edge and Cloud: An Early Exit based Clustering Approach
di: Bajpai, Divya Jyoti, et al.
Pubblicazione: (2024)
di: Bajpai, Divya Jyoti, et al.
Pubblicazione: (2024)
DistrEE: Distributed Early Exit of Deep Neural Network Inference on Edge Devices
di: Peng, Xian, et al.
Pubblicazione: (2025)
di: Peng, Xian, et al.
Pubblicazione: (2025)
Collaborative Speculative Inference for Efficient LLM Inference Serving
di: Gao, Luyao, et al.
Pubblicazione: (2025)
di: Gao, Luyao, et al.
Pubblicazione: (2025)
EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism
di: Chen, Yanxi, et al.
Pubblicazione: (2023)
di: Chen, Yanxi, et al.
Pubblicazione: (2023)
Adaptive Resolution Inference (ARI): Energy-Efficient Machine Learning for Internet of Things
di: Wang, Ziheng, et al.
Pubblicazione: (2024)
di: Wang, Ziheng, et al.
Pubblicazione: (2024)
Designing Large Foundation Models for Efficient Training and Inference: A Survey
di: Liu, Dong, et al.
Pubblicazione: (2024)
di: Liu, Dong, et al.
Pubblicazione: (2024)
Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection
di: Gupta, Vima, et al.
Pubblicazione: (2024)
di: Gupta, Vima, et al.
Pubblicazione: (2024)
InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management
di: Lee, Wonbeom, et al.
Pubblicazione: (2024)
di: Lee, Wonbeom, et al.
Pubblicazione: (2024)
TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval
di: Lin, Chien-Yu, et al.
Pubblicazione: (2025)
di: Lin, Chien-Yu, et al.
Pubblicazione: (2025)
DHP: Efficient Scaling of MLLM Training with Dynamic Hybrid Parallelism
di: Niu, Yifan, et al.
Pubblicazione: (2026)
di: Niu, Yifan, et al.
Pubblicazione: (2026)
Early-Exit meets Model-Distributed Inference at Edge Networks
di: Colocrese, Marco, et al.
Pubblicazione: (2024)
di: Colocrese, Marco, et al.
Pubblicazione: (2024)
Fast Distributed Inference Serving for Large Language Models
di: Wu, Bingyang, et al.
Pubblicazione: (2023)
di: Wu, Bingyang, et al.
Pubblicazione: (2023)
Arctic Inference with Shift Parallelism: Fast and Efficient Open Source Inference System for Enterprise AI
di: Rajbhandari, Samyam, et al.
Pubblicazione: (2025)
di: Rajbhandari, Samyam, et al.
Pubblicazione: (2025)
70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)
di: Zhang, Tianyi, et al.
Pubblicazione: (2025)
di: Zhang, Tianyi, et al.
Pubblicazione: (2025)
Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous Dequantization
di: Li, Jinhao, et al.
Pubblicazione: (2023)
di: Li, Jinhao, et al.
Pubblicazione: (2023)
Recurrent Early Exits for Federated Learning with Heterogeneous Clients
di: Lee, Royson, et al.
Pubblicazione: (2024)
di: Lee, Royson, et al.
Pubblicazione: (2024)
Practical Performance Guarantees for Pipelined DNN Inference
di: Archer, Aaron, et al.
Pubblicazione: (2023)
di: Archer, Aaron, et al.
Pubblicazione: (2023)
FDC: Fast KV Dimensionality Compression for Efficient LLM Inference
di: Zhang, Zeyu, et al.
Pubblicazione: (2024)
di: Zhang, Zeyu, et al.
Pubblicazione: (2024)
PecSched: Preemptive and Efficient Cluster Scheduling for LLM Inference
di: Zhang, Zeyu, et al.
Pubblicazione: (2024)
di: Zhang, Zeyu, et al.
Pubblicazione: (2024)
Kraken: Inherently Parallel Transformers For Efficient Multi-Device Inference
di: Prabhakar, Rohan Baskar, et al.
Pubblicazione: (2024)
di: Prabhakar, Rohan Baskar, et al.
Pubblicazione: (2024)
Split CNN Inference on Networked Microcontrollers
di: Lu, Junyu, et al.
Pubblicazione: (2026)
di: Lu, Junyu, et al.
Pubblicazione: (2026)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
di: Liu, Mengfan, et al.
Pubblicazione: (2025)
di: Liu, Mengfan, et al.
Pubblicazione: (2025)
Pie: Pooling CPU Memory for LLM Inference
di: Xu, Yi, et al.
Pubblicazione: (2024)
di: Xu, Yi, et al.
Pubblicazione: (2024)
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
di: Gond, Raja, et al.
Pubblicazione: (2025)
di: Gond, Raja, et al.
Pubblicazione: (2025)
Deal: Distributed End-to-End GNN Inference for All Nodes
di: Chen, Shiyang, et al.
Pubblicazione: (2025)
di: Chen, Shiyang, et al.
Pubblicazione: (2025)
Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions
di: Barrak, Amine, et al.
Pubblicazione: (2025)
di: Barrak, Amine, et al.
Pubblicazione: (2025)
PackInfer: Compute- and I/O-Efficient Attention for Batched LLM Inference
di: Ning, Rui, et al.
Pubblicazione: (2026)
di: Ning, Rui, et al.
Pubblicazione: (2026)
Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud
di: Ghosh, Himel
Pubblicazione: (2024)
di: Ghosh, Himel
Pubblicazione: (2024)
The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project
di: Chen, Huamin, et al.
Pubblicazione: (2026)
di: Chen, Huamin, et al.
Pubblicazione: (2026)
Quasar: Quantized Self-Speculative Acceleration for Rapid Inference via Memory-Efficient Verification
di: Huang, Guang, et al.
Pubblicazione: (2026)
di: Huang, Guang, et al.
Pubblicazione: (2026)
SuperFedNAS: Cost-Efficient Federated Neural Architecture Search for On-Device Inference
di: Khare, Alind, et al.
Pubblicazione: (2023)
di: Khare, Alind, et al.
Pubblicazione: (2023)
MobiZO: Enabling Efficient LLM Fine-Tuning at the Edge via Inference Engines
di: Gao, Lei, et al.
Pubblicazione: (2024)
di: Gao, Lei, et al.
Pubblicazione: (2024)
DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference
di: Zhang, Yujie, et al.
Pubblicazione: (2024)
di: Zhang, Yujie, et al.
Pubblicazione: (2024)
ChunkFlow: Communication-Aware Chunked Prefetching for Layerwise Offloading in Distributed Diffusion Transformer Inference
di: Meng, Han, et al.
Pubblicazione: (2026)
di: Meng, Han, et al.
Pubblicazione: (2026)
DALI: A Workload-Aware Offloading Framework for Efficient MoE Inference on Local PCs
di: Zhu, Zeyu, et al.
Pubblicazione: (2026)
di: Zhu, Zeyu, et al.
Pubblicazione: (2026)
Relay Buffer Independent Communication over Pooled HBM for Efficient MoE Inference on Ascend
di: Hu, Tianlun, et al.
Pubblicazione: (2026)
di: Hu, Tianlun, et al.
Pubblicazione: (2026)
Documenti analoghi
-
A Performance Analyzer for a Public Cloud's ML-Augmented VM Allocator
di: Bostandoost, Roozbeh, et al.
Pubblicazione: (2025) -
Towards Safer Heuristics With XPlain
di: Karimi, Pantea, et al.
Pubblicazione: (2024) -
Federated Learning for Collaborative Inference Systems: The Case of Early Exit Networks
di: Kaplan, Caelin, et al.
Pubblicazione: (2024) -
Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving
di: Dai, Yinwei, et al.
Pubblicazione: (2023) -
Distributed Inference on Mobile Edge and Cloud: An Early Exit based Clustering Approach
di: Bajpai, Divya Jyoti, et al.
Pubblicazione: (2024)