Saved in:
| Main Authors: | Wang, Kun, Cao, Jiani, Zhou, Zimu, Li, Zhenjiang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2401.16757 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hermes: Memory-Efficient Pipeline Inference for Large Models on Edge Devices
by: Han, Xueyuan, et al.
Published: (2024)
by: Han, Xueyuan, et al.
Published: (2024)
HiDP: Hierarchical DNN Partitioning for Distributed Inference on Heterogeneous Edge Platforms
by: Taufique, Zain, et al.
Published: (2024)
by: Taufique, Zain, et al.
Published: (2024)
NestQuant: Post-Training Integer-Nesting Quantization for On-Device DNN
by: Xie, Jianhang, et al.
Published: (2025)
by: Xie, Jianhang, et al.
Published: (2025)
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
by: Shen, Zheyu, et al.
Published: (2025)
by: Shen, Zheyu, et al.
Published: (2025)
Bitnet.cpp: Efficient Edge Inference for Ternary LLMs
by: Wang, Jinheng, et al.
Published: (2025)
by: Wang, Jinheng, et al.
Published: (2025)
DistrEE: Distributed Early Exit of Deep Neural Network Inference on Edge Devices
by: Peng, Xian, et al.
Published: (2025)
by: Peng, Xian, et al.
Published: (2025)
PIPO: Pipelined Offloading for Efficient Inference on Consumer Devices
by: Liu, Yangyijian, et al.
Published: (2025)
by: Liu, Yangyijian, et al.
Published: (2025)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
by: Jiang, Xuanlin, et al.
Published: (2024)
by: Jiang, Xuanlin, et al.
Published: (2024)
When Foresight Pruning Meets Zeroth-Order Optimization: Efficient Federated Learning for Low-Memory Devices
by: Zhang, Pengyu, et al.
Published: (2024)
by: Zhang, Pengyu, et al.
Published: (2024)
Adaptive Workload Distribution for Accuracy-aware DNN Inference on Collaborative Edge Platforms
by: Taufique, Zain, et al.
Published: (2023)
by: Taufique, Zain, et al.
Published: (2023)
SmartMem: Layout Transformation Elimination and Adaptation for Efficient DNN Execution on Mobile
by: Niu, Wei, et al.
Published: (2024)
by: Niu, Wei, et al.
Published: (2024)
Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU
by: Levine, Reese, et al.
Published: (2026)
by: Levine, Reese, et al.
Published: (2026)
EdgeRL: Reinforcement Learning-driven Deep Learning Model Inference Optimization at Edge
by: Mounesan, Motahare, et al.
Published: (2024)
by: Mounesan, Motahare, et al.
Published: (2024)
GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism
by: Jeon, Byungsoo, et al.
Published: (2024)
by: Jeon, Byungsoo, et al.
Published: (2024)
Inference Offloading for Cost-Sensitive Binary Classification at the Edge
by: Moothedath, Vishnu Narayanan, et al.
Published: (2025)
by: Moothedath, Vishnu Narayanan, et al.
Published: (2025)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
by: Cao, Shiyi, et al.
Published: (2024)
by: Cao, Shiyi, et al.
Published: (2024)
Stochastic Sparse Attention for Memory-Bound Inference
by: Lee, Kyle, et al.
Published: (2026)
by: Lee, Kyle, et al.
Published: (2026)
Beyond Inference: Performance Analysis of DNN Server Overheads for Computer Vision
by: AbouElhamayed, Ahmed F., et al.
Published: (2024)
by: AbouElhamayed, Ahmed F., et al.
Published: (2024)
A Survey on Collaborative DNN Inference for Edge Intelligence
by: Ren, Weiqing, et al.
Published: (2022)
by: Ren, Weiqing, et al.
Published: (2022)
SFPrompt: Communication-Efficient Split Federated Fine-Tuning for Large Pre-Trained Models over Resource-Limited Devices
by: Cao, Linxiao, et al.
Published: (2024)
by: Cao, Linxiao, et al.
Published: (2024)
P3SL: Personalized Privacy-Preserving Split Learning on Heterogeneous Edge Devices
by: Fan, Wei, et al.
Published: (2025)
by: Fan, Wei, et al.
Published: (2025)
Distributed Inference on Mobile Edge and Cloud: An Early Exit based Clustering Approach
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
Learning the Optimal Path and DNN Partition for Collaborative Edge Inference
by: Huang, Yin, et al.
Published: (2024)
by: Huang, Yin, et al.
Published: (2024)
FedMHO: Heterogeneous One-Shot Federated Learning Towards Resource-Constrained Edge Devices
by: Yao, Dezhong, et al.
Published: (2025)
by: Yao, Dezhong, et al.
Published: (2025)
MoEBlaze: Breaking the Memory Wall for Efficient MoE Training on Modern GPUs
by: Zhang, Jiyuan, et al.
Published: (2026)
by: Zhang, Jiyuan, et al.
Published: (2026)
Federated Attention: A Distributed Paradigm for Collaborative LLM Inference over Edge Networks
by: Deng, Xiumei, et al.
Published: (2025)
by: Deng, Xiumei, et al.
Published: (2025)
SparOA: Sparse and Operator-aware Hybrid Scheduling for Edge DNN Inference
by: Zhang, Ziyang, et al.
Published: (2025)
by: Zhang, Ziyang, et al.
Published: (2025)
ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling
by: Yang, Yuchen, et al.
Published: (2026)
by: Yang, Yuchen, et al.
Published: (2026)
Multi-Dimensional Autoscaling of Stream Processing Services on Edge Devices
by: Sedlak, Boris, et al.
Published: (2025)
by: Sedlak, Boris, et al.
Published: (2025)
Efficient Federated Finetuning of Tiny Transformers with Resource-Constrained Devices
by: Pfeiffer, Kilian, et al.
Published: (2024)
by: Pfeiffer, Kilian, et al.
Published: (2024)
Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer Inference
by: Ye, Shengyuan, et al.
Published: (2024)
by: Ye, Shengyuan, et al.
Published: (2024)
SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips
by: Yu, Jiahuan, et al.
Published: (2026)
by: Yu, Jiahuan, et al.
Published: (2026)
Online Client Scheduling and Resource Allocation for Efficient Federated Edge Learning
by: Gao, Zhidong, et al.
Published: (2024)
by: Gao, Zhidong, et al.
Published: (2024)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
by: Su, Zhaoyuan, et al.
Published: (2025)
by: Su, Zhaoyuan, et al.
Published: (2025)
MatKV: Trading Compute for Flash Storage in LLM Inference
by: Shin, Kun-Woo, et al.
Published: (2025)
by: Shin, Kun-Woo, et al.
Published: (2025)
Adaptive and Resource-efficient Agentic AI Systems for Mobile and Embedded Devices: A Survey
by: Liu, Sicong, et al.
Published: (2025)
by: Liu, Sicong, et al.
Published: (2025)
AgentStop: Terminating Local AI Agents Early to Save Energy in Consumer Devices
by: Pham, Dzung, et al.
Published: (2026)
by: Pham, Dzung, et al.
Published: (2026)
PGT-I: Scaling Spatiotemporal GNNs with Memory-Efficient Distributed Training
by: Ockerman, Seth, et al.
Published: (2025)
by: Ockerman, Seth, et al.
Published: (2025)
Two-Timescale Model Caching and Resource Allocation for Edge-Enabled AI-Generated Content Services
by: Liu, Zhang, et al.
Published: (2024)
by: Liu, Zhang, et al.
Published: (2024)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
by: Ye, Zihao, et al.
Published: (2025)
by: Ye, Zihao, et al.
Published: (2025)
Similar Items
-
Hermes: Memory-Efficient Pipeline Inference for Large Models on Edge Devices
by: Han, Xueyuan, et al.
Published: (2024) -
HiDP: Hierarchical DNN Partitioning for Distributed Inference on Heterogeneous Edge Platforms
by: Taufique, Zain, et al.
Published: (2024) -
NestQuant: Post-Training Integer-Nesting Quantization for On-Device DNN
by: Xie, Jianhang, et al.
Published: (2025) -
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
by: Shen, Zheyu, et al.
Published: (2025) -
Bitnet.cpp: Efficient Edge Inference for Ternary LLMs
by: Wang, Jinheng, et al.
Published: (2025)