Exploring the Dynamic Scheduling Space of Real-Time Generative AI Applications on Emerging Heterogeneous Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Karami, Rachid, Patwari, Rajeev, Kwon, Hyoukjun, Sirasao, Ashish |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Recover-LoRA: Data-Free Accuracy Recovery of Degraded Language Models via Low-Rank Adaptation
von: Das, Devleena, et al.
Veröffentlicht: (2025)
von: Das, Devleena, et al.
Veröffentlicht: (2025)
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
von: Patwari, Rajeev, et al.
Veröffentlicht: (2025)
von: Patwari, Rajeev, et al.
Veröffentlicht: (2025)
KV Pareto: Systems-Level Optimization of KV Cache and Model Compression for Long Context Inference
von: Gokhale, Sai, et al.
Veröffentlicht: (2025)
von: Gokhale, Sai, et al.
Veröffentlicht: (2025)
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
von: Karami, Rachid, et al.
Veröffentlicht: (2024)
von: Karami, Rachid, et al.
Veröffentlicht: (2024)
Characterizing State Space Model and Hybrid Language Model Performance with Long Context
von: Mitra, Saptarshi, et al.
Veröffentlicht: (2025)
von: Mitra, Saptarshi, et al.
Veröffentlicht: (2025)
SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module Accelerators
von: Odema, Mohanad, et al.
Veröffentlicht: (2024)
von: Odema, Mohanad, et al.
Veröffentlicht: (2024)
Efficient Depth Estimation for Unstable Stereo Camera Systems on AR Glasses
von: Liu, Yongfan, et al.
Veröffentlicht: (2024)
von: Liu, Yongfan, et al.
Veröffentlicht: (2024)
D-com: Accelerating Iterative Processing to Enable Low-rank Decomposition of Activations
von: Tahmasebi, Faraz, et al.
Veröffentlicht: (2025)
von: Tahmasebi, Faraz, et al.
Veröffentlicht: (2025)
Characterizing the Accuracy -- Efficiency Trade-off of Low-rank Decomposition in Language Models
von: Moar, Chakshu, et al.
Veröffentlicht: (2024)
von: Moar, Chakshu, et al.
Veröffentlicht: (2024)
Real-Time Evaluation Models for RAG: Who Detects Hallucinations Best?
von: Sardana, Ashish
Veröffentlicht: (2025)
von: Sardana, Ashish
Veröffentlicht: (2025)
FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression
von: Qiao, Ye, et al.
Veröffentlicht: (2026)
von: Qiao, Ye, et al.
Veröffentlicht: (2026)
HiGen: Hierarchical Graph Generative Networks
von: Karami, Mahdi
Veröffentlicht: (2023)
von: Karami, Mahdi
Veröffentlicht: (2023)
TimEHR: Image-based Time Series Generation for Electronic Health Records
von: Karami, Hojjat, et al.
Veröffentlicht: (2024)
von: Karami, Hojjat, et al.
Veröffentlicht: (2024)
A Comprehensive Forecasting-Based Framework for Time Series Anomaly Detection: Benchmarking on the Numenta Anomaly Benchmark (NAB)
von: Karami, Mohammad, et al.
Veröffentlicht: (2025)
von: Karami, Mohammad, et al.
Veröffentlicht: (2025)
Agile Reinforcement Learning for Real-Time Task Scheduling in Edge Computing
von: Avan, Amin, et al.
Veröffentlicht: (2025)
von: Avan, Amin, et al.
Veröffentlicht: (2025)
A Training-Time Diagnostic for Generalization via the Log-Alignment Ratio
von: Shehper, Ali, et al.
Veröffentlicht: (2026)
von: Shehper, Ali, et al.
Veröffentlicht: (2026)
Physics-Aware Heterogeneous GNN Architecture for Real-Time BESS Optimization in Unbalanced Distribution Systems
von: Ma, Aoxiang, et al.
Veröffentlicht: (2025)
von: Ma, Aoxiang, et al.
Veröffentlicht: (2025)
Dynamic Context Evolution for Scalable Synthetic Data Generation
von: Lingo, Ryan, et al.
Veröffentlicht: (2026)
von: Lingo, Ryan, et al.
Veröffentlicht: (2026)
Graph Anomaly Detection in Time Series: A Survey
von: Ho, Thi Kieu Khanh, et al.
Veröffentlicht: (2023)
von: Ho, Thi Kieu Khanh, et al.
Veröffentlicht: (2023)
Enhancing One-shot Pruned Pre-trained Language Models through Sparse-Dense-Sparse Mechanism
von: Li, Guanchen, et al.
Veröffentlicht: (2024)
von: Li, Guanchen, et al.
Veröffentlicht: (2024)
Time-Shifted Token Scheduling for Symbolic Music Generation
von: Wang, Ting-Kang, et al.
Veröffentlicht: (2025)
von: Wang, Ting-Kang, et al.
Veröffentlicht: (2025)
Dynamics of Learning: Generative Schedules from Latent ODEs
von: Sampson, Matt L., et al.
Veröffentlicht: (2025)
von: Sampson, Matt L., et al.
Veröffentlicht: (2025)
Design and Scheduling of an AI-based Queueing System
von: Lee, Jiung, et al.
Veröffentlicht: (2024)
von: Lee, Jiung, et al.
Veröffentlicht: (2024)
MS-SSM: A Multi-Scale State Space Model for Efficient Sequence Modeling
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
Auto-Regressive Masked Diffusion Models
von: Karami, Mahdi, et al.
Veröffentlicht: (2026)
von: Karami, Mahdi, et al.
Veröffentlicht: (2026)
Orchid: Flexible and Data-Dependent Convolution for Sequence Modeling
von: Karami, Mahdi, et al.
Veröffentlicht: (2024)
von: Karami, Mahdi, et al.
Veröffentlicht: (2024)
Exploring the Promise and Limits of Real-Time Recurrent Learning
von: Irie, Kazuki, et al.
Veröffentlicht: (2023)
von: Irie, Kazuki, et al.
Veröffentlicht: (2023)
A Finite Element-Inspired Hypergraph Neural Network: Application to Fluid Dynamics Simulations
von: Gao, Rui, et al.
Veröffentlicht: (2022)
von: Gao, Rui, et al.
Veröffentlicht: (2022)
Entropic Time Schedulers for Generative Diffusion Models
von: Stancevic, Dejan, et al.
Veröffentlicht: (2025)
von: Stancevic, Dejan, et al.
Veröffentlicht: (2025)
Robust Learning of Heterogeneous Dynamic Systems
von: Xu, Shuoxun, et al.
Veröffentlicht: (2026)
von: Xu, Shuoxun, et al.
Veröffentlicht: (2026)
DynamicBench: Evaluating Real-Time Report Generation in Large Language Models
von: Li, Jingyao, et al.
Veröffentlicht: (2025)
von: Li, Jingyao, et al.
Veröffentlicht: (2025)
Exploring the Potential of Synthetic Data to Replace Real Data
von: Lee, Hyungtae, et al.
Veröffentlicht: (2024)
von: Lee, Hyungtae, et al.
Veröffentlicht: (2024)
Generative Profiling for Soft Real-Time Systems and its Applications to Resource Allocation
von: Bondar, Georgiy A., et al.
Veröffentlicht: (2026)
von: Bondar, Georgiy A., et al.
Veröffentlicht: (2026)
CNN-Enabled Scheduling for Probabilistic Real-Time Guarantees in Industrial URLLC
von: Alqudah, Eman, et al.
Veröffentlicht: (2025)
von: Alqudah, Eman, et al.
Veröffentlicht: (2025)
FIRMA: FIbonacci Ring Model Aggregation for Privacy-preserving Federated Learning
von: Hedjam, Rachid
Veröffentlicht: (2026)
von: Hedjam, Rachid
Veröffentlicht: (2026)
Lévy-Flow Models: Heavy-Tail-Aware Normalizing Flows for Financial Risk Management
von: Drissi, Rachid
Veröffentlicht: (2026)
von: Drissi, Rachid
Veröffentlicht: (2026)
Exploring System Adaptations For Minimum Latency Real-Time Piano Transcription
von: Hu, Patricia, et al.
Veröffentlicht: (2025)
von: Hu, Patricia, et al.
Veröffentlicht: (2025)
A Real-Time Lyrics Alignment System Using Chroma And Phonetic Features For Classical Vocal Performance
von: Park, Jiyun, et al.
Veröffentlicht: (2024)
von: Park, Jiyun, et al.
Veröffentlicht: (2024)
Exploring Representation-Aligned Latent Space for Better Generation
von: Xu, Wanghan, et al.
Veröffentlicht: (2025)
von: Xu, Wanghan, et al.
Veröffentlicht: (2025)
Homogeneous Dynamics Space for Heterogeneous Humans
von: Liu, Xinpeng, et al.
Veröffentlicht: (2024)
von: Liu, Xinpeng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Recover-LoRA: Data-Free Accuracy Recovery of Degraded Language Models via Low-Rank Adaptation
von: Das, Devleena, et al.
Veröffentlicht: (2025) -
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
von: Patwari, Rajeev, et al.
Veröffentlicht: (2025) -
KV Pareto: Systems-Level Optimization of KV Cache and Model Compression for Long Context Inference
von: Gokhale, Sai, et al.
Veröffentlicht: (2025) -
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
von: Karami, Rachid, et al.
Veröffentlicht: (2024) -
Characterizing State Space Model and Hybrid Language Model Performance with Long Context
von: Mitra, Saptarshi, et al.
Veröffentlicht: (2025)