FusedInf: Efficient Swapping of DNN Models for On-Demand Serverless Inference Services on the Edge
Fuente:
arXiv
Saved in:
| Main Authors: | Taki, Sifat Ut, Padmanabhan, Arthi, Mastorakis, Spyridon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UnifiedNN: Efficient Neural Network Training on the Cloud
by: Taki, Sifat Ut, et al.
Published: (2024)
by: Taki, Sifat Ut, et al.
Published: (2024)
Amalgam: A Framework for Obfuscated Neural Network Training on the Cloud
by: Taki, Sifat Ut, et al.
Published: (2024)
by: Taki, Sifat Ut, et al.
Published: (2024)
SwapNet: Efficient Swapping for DNN Inference on Edge AI Devices Beyond the Memory Budget
by: Wang, Kun, et al.
Published: (2024)
by: Wang, Kun, et al.
Published: (2024)
AdaPI: Facilitating DNN Model Adaptivity for Efficient Private Inference in Edge Computing
by: Zhou, Tong, et al.
Published: (2024)
by: Zhou, Tong, et al.
Published: (2024)
Multi-DNN Inference of Sparse Models on Edge SoCs
by: Luo, Jiawei, et al.
Published: (2026)
by: Luo, Jiawei, et al.
Published: (2026)
Privacy-Aware Joint DNN Model Deployment and Partitioning Optimization for Collaborative Edge Inference Services
by: Cheng, Zhipeng, et al.
Published: (2025)
by: Cheng, Zhipeng, et al.
Published: (2025)
A Survey on Collaborative DNN Inference for Edge Intelligence
by: Ren, Weiqing, et al.
Published: (2022)
by: Ren, Weiqing, et al.
Published: (2022)
ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
by: Fu, Yao, et al.
Published: (2024)
by: Fu, Yao, et al.
Published: (2024)
EdgeFLow: Serverless Federated Learning via Sequential Model Migration in Edge Networks
by: Shi, Yuchen, et al.
Published: (2026)
by: Shi, Yuchen, et al.
Published: (2026)
Learning the Optimal Path and DNN Partition for Collaborative Edge Inference
by: Huang, Yin, et al.
Published: (2024)
by: Huang, Yin, et al.
Published: (2024)
THESAURUS: Contrastive Graph Clustering by Swapping Fused Gromov-Wasserstein Couplings
by: Deng, Bowen, et al.
Published: (2024)
by: Deng, Bowen, et al.
Published: (2024)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
by: Liu, Qunyou, et al.
Published: (2026)
by: Liu, Qunyou, et al.
Published: (2026)
NMS: Efficient Edge DNN Training via Near-Memory Sampling on Manifolds
by: Zhao, Boran, et al.
Published: (2025)
by: Zhao, Boran, et al.
Published: (2025)
Characterizing Encrypted Application Traffic through Cellular Radio Interface Protocol
by: Islam, Md Ruman, et al.
Published: (2024)
by: Islam, Md Ruman, et al.
Published: (2024)
Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud
by: Ghosh, Himel
Published: (2024)
by: Ghosh, Himel
Published: (2024)
InfAlign: Inference-aware language model alignment
by: Balashankar, Ananth, et al.
Published: (2024)
by: Balashankar, Ananth, et al.
Published: (2024)
DataInf: Efficiently Estimating Data Influence in LoRA-tuned LLMs and Diffusion Models
by: Kwon, Yongchan, et al.
Published: (2023)
by: Kwon, Yongchan, et al.
Published: (2023)
EdgeJury: Cross-Reviewed Small-Model Ensembles for Truthful Question Answering on Serverless Edge Inference
by: Kumar, Aayush
Published: (2025)
by: Kumar, Aayush
Published: (2025)
DNNShifter: An Efficient DNN Pruning System for Edge Computing
by: Eccles, Bailey J., et al.
Published: (2023)
by: Eccles, Bailey J., et al.
Published: (2023)
OpInf-LLM: Parametric PDE Solving with LLMs via Operator Inference
by: Wang, Zhuoyuan, et al.
Published: (2026)
by: Wang, Zhuoyuan, et al.
Published: (2026)
Efficient Swap Multicalibration of Elicitable Properties
by: Hu, Lunjia, et al.
Published: (2025)
by: Hu, Lunjia, et al.
Published: (2025)
Safe Multi-Agent Deep Reinforcement Learning for Privacy-Aware Edge-Device Collaborative DNN Inference
by: Wang, Hong, et al.
Published: (2026)
by: Wang, Hong, et al.
Published: (2026)
HiDP: Hierarchical DNN Partitioning for Distributed Inference on Heterogeneous Edge Platforms
by: Taufique, Zain, et al.
Published: (2024)
by: Taufique, Zain, et al.
Published: (2024)
SemShareKV: Efficient KVCache Sharing for Semantically Similar Prompts via Token-Level LSH Matching
by: Zhao, Xinye, et al.
Published: (2025)
by: Zhao, Xinye, et al.
Published: (2025)
Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions
by: Barrak, Amine, et al.
Published: (2025)
by: Barrak, Amine, et al.
Published: (2025)
Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU
by: Ning, Zhenyu, et al.
Published: (2024)
by: Ning, Zhenyu, et al.
Published: (2024)
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
by: Sui, Yifan, et al.
Published: (2025)
by: Sui, Yifan, et al.
Published: (2025)
Efficient Swap Regret Minimization in Combinatorial Bandits
by: Kontogiannis, Andreas, et al.
Published: (2026)
by: Kontogiannis, Andreas, et al.
Published: (2026)
Hardware-Aware DNN Compression for Homogeneous Edge Devices
by: Zhang, Kunlong, et al.
Published: (2025)
by: Zhang, Kunlong, et al.
Published: (2025)
Leveraging Highly Approximated Multipliers in DNN Inference
by: Zervakis, Georgios, et al.
Published: (2024)
by: Zervakis, Georgios, et al.
Published: (2024)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
by: Liu, Mengfan, et al.
Published: (2025)
by: Liu, Mengfan, et al.
Published: (2025)
Deduplicator: When Computation Reuse Meets Load Balancing at the Network Edge
by: Azad, Md Washik Al, et al.
Published: (2024)
by: Azad, Md Washik Al, et al.
Published: (2024)
Improved Bounds for Swap Multicalibration and Swap Omniprediction
by: Luo, Haipeng, et al.
Published: (2025)
by: Luo, Haipeng, et al.
Published: (2025)
Adaptive Workload Distribution for Accuracy-aware DNN Inference on Collaborative Edge Platforms
by: Taufique, Zain, et al.
Published: (2023)
by: Taufique, Zain, et al.
Published: (2023)
The Larger the Merrier? Efficient Large AI Model Inference in Wireless Edge Networks
by: Lyu, Zhonghao, et al.
Published: (2025)
by: Lyu, Zhonghao, et al.
Published: (2025)
Inf2Guard: An Information-Theoretic Framework for Learning Privacy-Preserving Representations against Inference Attacks
by: Noorbakhsh, Sayedeh Leila, et al.
Published: (2024)
by: Noorbakhsh, Sayedeh Leila, et al.
Published: (2024)
Optimizing DNN Inference on Multi-Accelerator SoCs at Training-time
by: Risso, Matteo, et al.
Published: (2024)
by: Risso, Matteo, et al.
Published: (2024)
DVFS-Aware DNN Inference on GPUs: Latency Modeling and Performance Analysis
by: Han, Yunchu, et al.
Published: (2025)
by: Han, Yunchu, et al.
Published: (2025)
Memory-Efficient Partitioned DNN Inference on Resource-Constrained Android Crowds
by: Manamperi, Lakshani, et al.
Published: (2026)
by: Manamperi, Lakshani, et al.
Published: (2026)
FairyFuse: Multiplication-Free LLM Inference on CPUs via Fused Ternary Kernels
by: Zuo, Fei, et al.
Published: (2026)
by: Zuo, Fei, et al.
Published: (2026)
Similar Items
-
UnifiedNN: Efficient Neural Network Training on the Cloud
by: Taki, Sifat Ut, et al.
Published: (2024) -
Amalgam: A Framework for Obfuscated Neural Network Training on the Cloud
by: Taki, Sifat Ut, et al.
Published: (2024) -
SwapNet: Efficient Swapping for DNN Inference on Edge AI Devices Beyond the Memory Budget
by: Wang, Kun, et al.
Published: (2024) -
AdaPI: Facilitating DNN Model Adaptivity for Efficient Private Inference in Edge Computing
by: Zhou, Tong, et al.
Published: (2024) -
Multi-DNN Inference of Sparse Models on Edge SoCs
by: Luo, Jiawei, et al.
Published: (2026)