Puzzle: Scheduling Multiple Deep Learning Models on Mobile Device with Heterogeneous Processors
Fuente:
arXiv
Saved in:
| Main Authors: | Kang, Duseok, Lee, Yunseong, Kim, Junghoon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TempoNet: Slack-Quantized Transformer-Guided Reinforcement Scheduler for Adaptive Deadline-Centric Real-Time Dispatchs
by: Fu, Rong, et al.
Published: (2026)
by: Fu, Rong, et al.
Published: (2026)
Accelerated Training on Low-Power Edge Devices
by: Ahmed, Mohamed Aboelenien, et al.
Published: (2025)
by: Ahmed, Mohamed Aboelenien, et al.
Published: (2025)
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
by: Wu, Yinpeng, et al.
Published: (2026)
by: Wu, Yinpeng, et al.
Published: (2026)
MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems
by: Wang, Yifei, et al.
Published: (2026)
by: Wang, Yifei, et al.
Published: (2026)
Energy-Efficient Computation with DVFS using Deep Reinforcement Learning for Multi-Task Systems in Edge Computing
by: Li, Xinyi, et al.
Published: (2024)
by: Li, Xinyi, et al.
Published: (2024)
Semantic Scheduling for LLM Inference
by: Hua, Wenyue, et al.
Published: (2025)
by: Hua, Wenyue, et al.
Published: (2025)
Reinforcement Learning for Dynamic Memory Allocation
by: Lim, Arisrei, et al.
Published: (2024)
by: Lim, Arisrei, et al.
Published: (2024)
Machine Learning (ML) library in Linux kernel
by: Dubeyko, Viacheslav
Published: (2026)
by: Dubeyko, Viacheslav
Published: (2026)
LithOS: An Operating System for Efficient Machine Learning on GPUs
by: Coppock, Patrick H., et al.
Published: (2025)
by: Coppock, Patrick H., et al.
Published: (2025)
MaLV-OS: Rethinking the Operating System Architecture for Machine Learning in Virtualized Clouds
by: Bitchebe, Stella, et al.
Published: (2025)
by: Bitchebe, Stella, et al.
Published: (2025)
Beyond Edge Coverage: Per-Task Data-Flow Extraction at Kernel Function Boundaries via LLVM
by: Kim, Yunseong
Published: (2026)
by: Kim, Yunseong
Published: (2026)
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
by: Song, Yixin, et al.
Published: (2023)
by: Song, Yixin, et al.
Published: (2023)
Enhancing Battery Storage Energy Arbitrage with Deep Reinforcement Learning and Time-Series Forecasting
by: Sage, Manuel, et al.
Published: (2024)
by: Sage, Manuel, et al.
Published: (2024)
E-Mapper: Energy-Efficient Resource Allocation for Traditional Operating Systems on Heterogeneous Processors
by: Smejkal, Till, et al.
Published: (2024)
by: Smejkal, Till, et al.
Published: (2024)
Enhancing Adaptive Mixed-Criticality Scheduling with Deep Reinforcement Learning
by: Mendes, Bruno, et al.
Published: (2024)
by: Mendes, Bruno, et al.
Published: (2024)
Crash-Consistent Checkpointing for AI Training on macOS/APFS
by: Jeon, Juha
Published: (2025)
by: Jeon, Juha
Published: (2025)
Herding LLaMaS: Using LLMs as an OS Module
by: Kamath, Aditya K, et al.
Published: (2024)
by: Kamath, Aditya K, et al.
Published: (2024)
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
by: Prabhu, Ramya, et al.
Published: (2024)
by: Prabhu, Ramya, et al.
Published: (2024)
LLM as a System Service on Mobile Devices
by: Yin, Wangsong, et al.
Published: (2024)
by: Yin, Wangsong, et al.
Published: (2024)
KernelOracle: Predicting the Linux Scheduler's Next Move with Deep Learning
by: Kahu, Sampanna Yashwant
Published: (2025)
by: Kahu, Sampanna Yashwant
Published: (2025)
Leveraging Machine Learning for Accurate IoT Device Identification in Dynamic Wireless Contexts
by: Tushir, Bhagyashri, et al.
Published: (2024)
by: Tushir, Bhagyashri, et al.
Published: (2024)
Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With Seneca
by: Desai, Omkar, et al.
Published: (2025)
by: Desai, Omkar, et al.
Published: (2025)
When eBPF Meets Machine Learning: On-the-fly OS Kernel Compartmentalization
by: Wang, Zicheng, et al.
Published: (2024)
by: Wang, Zicheng, et al.
Published: (2024)
Holistic Heterogeneous Scheduling for Autonomous Applications using Fine-grained, Multi-XPU Abstraction
by: Han, Mingcong, et al.
Published: (2025)
by: Han, Mingcong, et al.
Published: (2025)
Bauplan: zero-copy, scale-up FaaS for data pipelines
by: Tagliabue, Jacopo, et al.
Published: (2024)
by: Tagliabue, Jacopo, et al.
Published: (2024)
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
Modeling and Scheduling of Fusion Patterns in Autonomous Driving Systems (Extended Version)
by: Sobhani, Hoora, et al.
Published: (2025)
by: Sobhani, Hoora, et al.
Published: (2025)
OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents
by: Abhyankar, Reyna, et al.
Published: (2025)
by: Abhyankar, Reyna, et al.
Published: (2025)
Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
by: Chu, Kexin, et al.
Published: (2025)
by: Chu, Kexin, et al.
Published: (2025)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
From Imperative to Declarative: Towards LLM-friendly OS Interfaces for Boosted Computer-Use Agents
by: Wang, Yuan, et al.
Published: (2025)
by: Wang, Yuan, et al.
Published: (2025)
An Integrated Artificial Intelligence Operating System for Advanced Low-Altitude Aviation Applications
by: Tan, Minzhe, et al.
Published: (2024)
by: Tan, Minzhe, et al.
Published: (2024)
ConsumerBench: Benchmarking Generative AI Applications on End-User Devices
by: Gu, Yile, et al.
Published: (2025)
by: Gu, Yile, et al.
Published: (2025)
Generative Profiling for Soft Real-Time Systems and its Applications to Resource Allocation
by: Bondar, Georgiy A., et al.
Published: (2026)
by: Bondar, Georgiy A., et al.
Published: (2026)
Qurator: Scheduling Hybrid Quantum-Classical Workflows Across Heterogeneous Cloud Providers
by: Pehlivanoglu, Sinan, et al.
Published: (2026)
by: Pehlivanoglu, Sinan, et al.
Published: (2026)
Learning Semantics, Not Addresses: Runtime Neural Prefetching for Far Memory
by: Huang, Yutong, et al.
Published: (2025)
by: Huang, Yutong, et al.
Published: (2025)
Dynamic Optimization of Storage Systems Using Reinforcement Learning Techniques
by: Cheng, Chiyu, et al.
Published: (2024)
by: Cheng, Chiyu, et al.
Published: (2024)
Dynamic Adaptation in Data Storage: Real-Time Machine Learning for Enhanced Prefetching
by: Cheng, Chiyu, et al.
Published: (2024)
by: Cheng, Chiyu, et al.
Published: (2024)
Optimizing SSD Caches for Cloud Block Storage Systems Using Machine Learning Approaches
by: Cheng, Chiyu, et al.
Published: (2024)
by: Cheng, Chiyu, et al.
Published: (2024)
Work-in-Progress: Multi-Deadline DAG Scheduling Model for Autonomous Driving Systems
by: Yano, Atsushi, et al.
Published: (2025)
by: Yano, Atsushi, et al.
Published: (2025)
Similar Items
-
TempoNet: Slack-Quantized Transformer-Guided Reinforcement Scheduler for Adaptive Deadline-Centric Real-Time Dispatchs
by: Fu, Rong, et al.
Published: (2026) -
Accelerated Training on Low-Power Edge Devices
by: Ahmed, Mohamed Aboelenien, et al.
Published: (2025) -
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
by: Wu, Yinpeng, et al.
Published: (2026) -
MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems
by: Wang, Yifei, et al.
Published: (2026) -
Energy-Efficient Computation with DVFS using Deep Reinforcement Learning for Multi-Task Systems in Edge Computing
by: Li, Xinyi, et al.
Published: (2024)