Reconstruction-Based Adaptive Scheduling Using AI Inferences in Safety-Critical Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Alshaer, Samer, Khalifeh, Ala, Obermaisser, Roman |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Adaptive Approach to Enhance Machine Learning Scheduling Algorithms During Runtime Using Reinforcement Learning in Metascheduling Applications
di: Alshaer, Samer, et al.
Pubblicazione: (2025)
di: Alshaer, Samer, et al.
Pubblicazione: (2025)
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
di: Wu, Qi, et al.
Pubblicazione: (2026)
di: Wu, Qi, et al.
Pubblicazione: (2026)
Mixture-of-Schedulers: An Adaptive Scheduling Agent as a Learned Router for Expert Policies
di: Wang, Xinbo, et al.
Pubblicazione: (2025)
di: Wang, Xinbo, et al.
Pubblicazione: (2025)
Keep Your Friends Close: Leveraging Affinity Groups to Accelerate AI Inference Workflows
di: Garrett, Thiago, et al.
Pubblicazione: (2023)
di: Garrett, Thiago, et al.
Pubblicazione: (2023)
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
di: Shen, Zixu, et al.
Pubblicazione: (2025)
di: Shen, Zixu, et al.
Pubblicazione: (2025)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
di: Pan, Xinglin, et al.
Pubblicazione: (2025)
di: Pan, Xinglin, et al.
Pubblicazione: (2025)
SparOA: Sparse and Operator-aware Hybrid Scheduling for Edge DNN Inference
di: Zhang, Ziyang, et al.
Pubblicazione: (2025)
di: Zhang, Ziyang, et al.
Pubblicazione: (2025)
TCM-Serve: Modality-aware Scheduling for Multimodal Large Language Model Inference
di: Papaioannou, Konstantinos, et al.
Pubblicazione: (2026)
di: Papaioannou, Konstantinos, et al.
Pubblicazione: (2026)
Compass: A Decentralized Scheduler for Latency-Sensitive ML Workflows
di: Yang, Yuting, et al.
Pubblicazione: (2024)
di: Yang, Yuting, et al.
Pubblicazione: (2024)
Autonomous Systems Dependability in the era of AI: Design Challenges in Safety, Security, Reliability and Certification
di: Ranjbar, Behnaz, et al.
Pubblicazione: (2026)
di: Ranjbar, Behnaz, et al.
Pubblicazione: (2026)
LeMix: Unified Scheduling for LLM Training and Inference on Multi-GPU Systems
di: Li, Yufei, et al.
Pubblicazione: (2025)
di: Li, Yufei, et al.
Pubblicazione: (2025)
A-IO: Adaptive Inference Orchestration for Memory-Bound NPUs
di: Zhang, Chen, et al.
Pubblicazione: (2026)
di: Zhang, Chen, et al.
Pubblicazione: (2026)
Evaluating the Efficacy of LLM-Based Reasoning for Multiobjective HPC Job Scheduling
di: Jadhav, Prachi, et al.
Pubblicazione: (2025)
di: Jadhav, Prachi, et al.
Pubblicazione: (2025)
MSCCL++: Rethinking GPU Communication Abstractions for AI Inference
di: Hwang, Changho, et al.
Pubblicazione: (2025)
di: Hwang, Changho, et al.
Pubblicazione: (2025)
Decentralized AI: Permissionless LLM Inference on POKT Network
di: Olshansky, Daniel, et al.
Pubblicazione: (2024)
di: Olshansky, Daniel, et al.
Pubblicazione: (2024)
FIRST: Federated Inference Resource Scheduling Toolkit for Scientific AI Model Access
di: Tanikanti, Aditya, et al.
Pubblicazione: (2025)
di: Tanikanti, Aditya, et al.
Pubblicazione: (2025)
Profiling-Driven Adaptive Distributed Transformer Inference on Embedded Edge Deployment
di: Qazi, Muhammad Azlan, et al.
Pubblicazione: (2026)
di: Qazi, Muhammad Azlan, et al.
Pubblicazione: (2026)
Automated Road Safety: Enhancing Sign and Surface Damage Detection with AI
di: Merolla, Davide, et al.
Pubblicazione: (2024)
di: Merolla, Davide, et al.
Pubblicazione: (2024)
TS-EoH: An Edge Server Task Scheduling Algorithm Based on Evolution of Heuristic
di: Yatong, Wang, et al.
Pubblicazione: (2024)
di: Yatong, Wang, et al.
Pubblicazione: (2024)
Duration-Informed Workload Scheduler
di: Loreti, Daniela, et al.
Pubblicazione: (2026)
di: Loreti, Daniela, et al.
Pubblicazione: (2026)
Workload Schedulers -- Genesis, Algorithms and Differences
di: Sliwko, Leszek, et al.
Pubblicazione: (2025)
di: Sliwko, Leszek, et al.
Pubblicazione: (2025)
Decentralized Distributed Proximal Policy Optimization (DD-PPO) for High Performance Computing Scheduling on Multi-User Systems
di: Sgambati, Matthew, et al.
Pubblicazione: (2025)
di: Sgambati, Matthew, et al.
Pubblicazione: (2025)
AI Inference as Relocatable Electricity Demand: A Latency-Constrained Energy-Geography Framework
di: Luo, Xubin, et al.
Pubblicazione: (2026)
di: Luo, Xubin, et al.
Pubblicazione: (2026)
Adaptive AI-based Decentralized Resource Management in the Cloud-Edge Continuum
di: Li, Lanpei, et al.
Pubblicazione: (2025)
di: Li, Lanpei, et al.
Pubblicazione: (2025)
ECCENTRIC: Edge-Cloud Collaboration Framework for Distributed Inference Using Knowledge Adaptation
di: Kamani, Mohammad Mahdi, et al.
Pubblicazione: (2025)
di: Kamani, Mohammad Mahdi, et al.
Pubblicazione: (2025)
Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks
di: Chandrasekar, Ashok, et al.
Pubblicazione: (2026)
di: Chandrasekar, Ashok, et al.
Pubblicazione: (2026)
Accelerating Latency-Critical Applications with AI-Powered Semi-Automatic Fine-Grained Parallelization on SMT Processors
di: Los, Denis, et al.
Pubblicazione: (2025)
di: Los, Denis, et al.
Pubblicazione: (2025)
Power- and Fragmentation-aware Online Scheduling for GPU Datacenters
di: Lettich, Francesco, et al.
Pubblicazione: (2024)
di: Lettich, Francesco, et al.
Pubblicazione: (2024)
Dynamic Scheduling Strategies for Resource Optimization in Computing Environments
di: Wang, Xiaoye
Pubblicazione: (2024)
di: Wang, Xiaoye
Pubblicazione: (2024)
Cloud-Based AI Systems: Leveraging Large Language Models for Intelligent Fault Detection and Autonomous Self-Healing
di: Ji, Cheng, et al.
Pubblicazione: (2025)
di: Ji, Cheng, et al.
Pubblicazione: (2025)
Towards Resource-Efficient Compound AI Systems
di: Chaudhry, Gohar Irfan, et al.
Pubblicazione: (2025)
di: Chaudhry, Gohar Irfan, et al.
Pubblicazione: (2025)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
di: Bai, Fan, et al.
Pubblicazione: (2026)
di: Bai, Fan, et al.
Pubblicazione: (2026)
Equinox: Holistic Fair Scheduling in Serving Large Language Models
di: Wei, Zhixiang, et al.
Pubblicazione: (2025)
di: Wei, Zhixiang, et al.
Pubblicazione: (2025)
Capacity Planning and Scheduling for Jobs with Uncertainty in Resource Usage and Duration
di: Patra, Sunandita, et al.
Pubblicazione: (2025)
di: Patra, Sunandita, et al.
Pubblicazione: (2025)
Topology-aware Preemptive Scheduling for Co-located LLM Workloads
di: Zhang, Ping, et al.
Pubblicazione: (2024)
di: Zhang, Ping, et al.
Pubblicazione: (2024)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
di: Yu, Dianhai, et al.
Pubblicazione: (2022)
di: Yu, Dianhai, et al.
Pubblicazione: (2022)
Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
di: Da, Wei, et al.
Pubblicazione: (2025)
di: Da, Wei, et al.
Pubblicazione: (2025)
HiveMind: OS-Inspired Scheduling for Concurrent LLM Agent Workloads
di: Agyemang, Justice Owusu, et al.
Pubblicazione: (2026)
di: Agyemang, Justice Owusu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Adaptive Approach to Enhance Machine Learning Scheduling Algorithms During Runtime Using Reinforcement Learning in Metascheduling Applications
di: Alshaer, Samer, et al.
Pubblicazione: (2025) -
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
di: Wu, Qi, et al.
Pubblicazione: (2026) -
Mixture-of-Schedulers: An Adaptive Scheduling Agent as a Learned Router for Expert Policies
di: Wang, Xinbo, et al.
Pubblicazione: (2025) -
Keep Your Friends Close: Leveraging Affinity Groups to Accelerate AI Inference Workflows
di: Garrett, Thiago, et al.
Pubblicazione: (2023) -
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
di: Shen, Zixu, et al.
Pubblicazione: (2025)