AutoTailor: Automatic and Efficient Adaptive Model Deployment for Diverse Edge Devices
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Mengyang, Lu, Chenyu, Tian, Haodong, Dong, Fang, Zhou, Ruiting, Wang, Wei, Shen, Dian, Li, Guangtong, Wan, Ye, Li, Li |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Joint Optimization of DNN Model Caching and Request Routing in Mobile Edge Computing
von: Qiu, Shuting, et al.
Veröffentlicht: (2025)
von: Qiu, Shuting, et al.
Veröffentlicht: (2025)
On-Demand Multi-Task Sparsity for Efficient Large-Model Deployment on Edge Devices
von: Huang, Lianming, et al.
Veröffentlicht: (2025)
von: Huang, Lianming, et al.
Veröffentlicht: (2025)
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
von: Shen, Zheyu, et al.
Veröffentlicht: (2025)
von: Shen, Zheyu, et al.
Veröffentlicht: (2025)
AutoPEFT: Automatic Configuration Search for Parameter-Efficient Fine-Tuning
von: Zhou, Han, et al.
Veröffentlicht: (2023)
von: Zhou, Han, et al.
Veröffentlicht: (2023)
Deploy DINO with Many-to-Many Association
von: Jiang, Haodong, et al.
Veröffentlicht: (2026)
von: Jiang, Haodong, et al.
Veröffentlicht: (2026)
Adaptive Device-Edge Collaboration on DNN Inference in AIoT: A Digital Twin-Assisted Approach
von: Hu, Shisheng, et al.
Veröffentlicht: (2024)
von: Hu, Shisheng, et al.
Veröffentlicht: (2024)
AutoBreach: Universal and Adaptive Jailbreaking with Efficient Wordplay-Guided Optimization
von: Chen, Jiawei, et al.
Veröffentlicht: (2024)
von: Chen, Jiawei, et al.
Veröffentlicht: (2024)
Empirical Guidelines for Deploying LLMs onto Resource-constrained Edge Devices
von: Qin, Ruiyang, et al.
Veröffentlicht: (2024)
von: Qin, Ruiyang, et al.
Veröffentlicht: (2024)
AyE-Edge: Automated Deployment Space Search Empowering Accuracy yet Efficient Real-Time Object Detection on the Edge
von: Wu, Chao, et al.
Veröffentlicht: (2024)
von: Wu, Chao, et al.
Veröffentlicht: (2024)
Model of Dark Current in Silicon‐Based Barrier Impurity Band Infrared Detector Devices
von: Mengyang Cui, et al.
Veröffentlicht: (2026)
von: Mengyang Cui, et al.
Veröffentlicht: (2026)
TMLC-Net: Transferable Meta Label Correction for Noisy Label Learning
von: Li, Mengyang
Veröffentlicht: (2025)
von: Li, Mengyang
Veröffentlicht: (2025)
AutoTriton: Automatic Triton Programming with Reinforcement Learning in LLMs
von: Li, Shangzhan, et al.
Veröffentlicht: (2025)
von: Li, Shangzhan, et al.
Veröffentlicht: (2025)
AutoQD: Automatic Discovery of Diverse Behaviors with Quality-Diversity Optimization
von: Hedayatian, Saeed, et al.
Veröffentlicht: (2025)
von: Hedayatian, Saeed, et al.
Veröffentlicht: (2025)
PI-Whisper: Designing an Adaptive and Incremental Automatic Speech Recognition System for Edge Devices
von: Nassereldine, Amir, et al.
Veröffentlicht: (2024)
von: Nassereldine, Amir, et al.
Veröffentlicht: (2024)
TinyFormer: Efficient Transformer Design and Deployment on Tiny Devices
von: Yang, Jianlei, et al.
Veröffentlicht: (2023)
von: Yang, Jianlei, et al.
Veröffentlicht: (2023)
TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment
von: Li, Jiaming, et al.
Veröffentlicht: (2026)
von: Li, Jiaming, et al.
Veröffentlicht: (2026)
AutoHete: An Automatic and Efficient Heterogeneous Training System for LLMs
von: Zeng, Zihao, et al.
Veröffentlicht: (2025)
von: Zeng, Zihao, et al.
Veröffentlicht: (2025)
Efficient Partitioning Vision Transformer on Edge Devices for Distributed Inference
von: Liu, Xiang, et al.
Veröffentlicht: (2024)
von: Liu, Xiang, et al.
Veröffentlicht: (2024)
Auto-PRE: An Automatic and Cost-Efficient Peer-Review Framework for Language Generation Evaluation
von: Chen, Junjie, et al.
Veröffentlicht: (2024)
von: Chen, Junjie, et al.
Veröffentlicht: (2024)
SwapNet: Efficient Swapping for DNN Inference on Edge AI Devices Beyond the Memory Budget
von: Wang, Kun, et al.
Veröffentlicht: (2024)
von: Wang, Kun, et al.
Veröffentlicht: (2024)
CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-Device
von: Lin, Ye, et al.
Veröffentlicht: (2026)
von: Lin, Ye, et al.
Veröffentlicht: (2026)
Slimmable ConvNeXt: Width-Adaptive Inference for Efficient Multi-Device Deployment
von: Haberer, Janek, et al.
Veröffentlicht: (2026)
von: Haberer, Janek, et al.
Veröffentlicht: (2026)
EdgeOL: Efficient in-situ Online Learning on Edge Devices
von: Li, Sheng, et al.
Veröffentlicht: (2024)
von: Li, Sheng, et al.
Veröffentlicht: (2024)
HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
von: Dong, Jiangwen, et al.
Veröffentlicht: (2025)
von: Dong, Jiangwen, et al.
Veröffentlicht: (2025)
Graph Neural Networks Automated Design and Deployment on Device-Edge Co-Inference Systems
von: Zhou, Ao, et al.
Veröffentlicht: (2024)
von: Zhou, Ao, et al.
Veröffentlicht: (2024)
DSSD: Efficient Edge-Device LLM Deployment and Collaborative Inference via Distributed Split Speculative Decoding
von: Ning, Jiahong, et al.
Veröffentlicht: (2025)
von: Ning, Jiahong, et al.
Veröffentlicht: (2025)
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
von: Liu, Wentao, et al.
Veröffentlicht: (2025)
von: Liu, Wentao, et al.
Veröffentlicht: (2025)
The MoE-Empowered Edge LLMs Deployment: Architecture, Challenges, and Opportunities
von: Li, Ning, et al.
Veröffentlicht: (2025)
von: Li, Ning, et al.
Veröffentlicht: (2025)
EdgeInfinite: A Memory-Efficient Infinite-Context Transformer for Edge Devices
von: Chen, Jiyu, et al.
Veröffentlicht: (2025)
von: Chen, Jiyu, et al.
Veröffentlicht: (2025)
FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning
von: Chen, Haodong, et al.
Veröffentlicht: (2025)
von: Chen, Haodong, et al.
Veröffentlicht: (2025)
AMED: Automatic Mixed-Precision Quantization for Edge Devices
von: Kimhi, Moshe, et al.
Veröffentlicht: (2022)
von: Kimhi, Moshe, et al.
Veröffentlicht: (2022)
AutoTour: Automatic Photo Tour Guide with Smartphones and LLMs
von: Xu, Huatao, et al.
Veröffentlicht: (2026)
von: Xu, Huatao, et al.
Veröffentlicht: (2026)
SlimEdge: Performance and Device Aware Distributed DNN Deployment on Resource-Constrained Edge Hardware
von: Kumar, Mahadev Sunil, et al.
Veröffentlicht: (2025)
von: Kumar, Mahadev Sunil, et al.
Veröffentlicht: (2025)
Towards Efficient Image Deblurring for Edge Deployment
von: Miriyala, Srinivas, et al.
Veröffentlicht: (2026)
von: Miriyala, Srinivas, et al.
Veröffentlicht: (2026)
Hunting the Ghost: Towards Automatic Mining of IoT Hidden Services
von: Dong, Shuaike, et al.
Veröffentlicht: (2025)
von: Dong, Shuaike, et al.
Veröffentlicht: (2025)
Tailoring Molecular Diffusion in Core‐Shell Zeolite Imidazolate Framework Composites Realizes Efficient Kinetic Separation of Xylene Isomers
von: Linghe Yang, et al.
Veröffentlicht: (2025)
von: Linghe Yang, et al.
Veröffentlicht: (2025)
Adaptive composite event‐triggered control of a deferred constrained nonlinear system
von: Chenyu Zhang, et al.
Veröffentlicht: (2025)
von: Chenyu Zhang, et al.
Veröffentlicht: (2025)
Efficient Cloud-Edge-Device Query Execution Based on Collaborative Scan Operator
von: Zhao, Chunyu, et al.
Veröffentlicht: (2025)
von: Zhao, Chunyu, et al.
Veröffentlicht: (2025)
Edge Learning Based Collaborative Automatic Modulation Classification for Hierarchical Cognitive Radio Networks
von: Dong, Peihao, et al.
Veröffentlicht: (2024)
von: Dong, Peihao, et al.
Veröffentlicht: (2024)
AutoMMLab: Automatically Generating Deployable Models from Language Instructions for Computer Vision Tasks
von: Yang, Zekang, et al.
Veröffentlicht: (2024)
von: Yang, Zekang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Joint Optimization of DNN Model Caching and Request Routing in Mobile Edge Computing
von: Qiu, Shuting, et al.
Veröffentlicht: (2025) -
On-Demand Multi-Task Sparsity for Efficient Large-Model Deployment on Edge Devices
von: Huang, Lianming, et al.
Veröffentlicht: (2025) -
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
von: Shen, Zheyu, et al.
Veröffentlicht: (2025) -
AutoPEFT: Automatic Configuration Search for Parameter-Efficient Fine-Tuning
von: Zhou, Han, et al.
Veröffentlicht: (2023) -
Deploy DINO with Many-to-Many Association
von: Jiang, Haodong, et al.
Veröffentlicht: (2026)