Adaptive Orchestration for Large-Scale Inference on Heterogeneous Accelerator Systems Balancing Cost, Performance, and Resilience
Fuente:
arXiv
Saved in:
| Main Authors: | Biran, Yahav, Kissos, Imry |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Orchestration for Large‐Scale Inference on Heterogeneous Accelerator Systems: Balancing Cost, Performance, and Resilience
by: Yahav Biran, et al.
Published: (2026)
by: Yahav Biran, et al.
Published: (2026)
Task-parallelism in SWIFT for heterogeneous compute architectures
by: Nasar, Abouzied M. A., et al.
Published: (2025)
by: Nasar, Abouzied M. A., et al.
Published: (2025)
Understanding Knowledge Transferability for Transfer Learning: A Survey
by: Wang, Haohua, et al.
Published: (2025)
by: Wang, Haohua, et al.
Published: (2025)
Deep Learning Model Deployment in Multiple Cloud Providers: an Exploratory Study Using Low Computing Power Environments
by: Lemos, Elayne, et al.
Published: (2025)
by: Lemos, Elayne, et al.
Published: (2025)
Systems Engineering of Adaptive AI Inference Orchestration Across Heterogeneous Accelerators
by: Yahav Biran
Published: (2026)
by: Yahav Biran
Published: (2026)
SPEC CPU: The Next Generation
by: Madhav, Mahesh, et al.
Published: (2026)
by: Madhav, Mahesh, et al.
Published: (2026)
MH-FSF: A Unified Framework for Overcoming Benchmarking and Reproducibility Limitations in Feature Selection Evaluation
by: Rocha, Vanderson, et al.
Published: (2025)
by: Rocha, Vanderson, et al.
Published: (2025)
Graph Transformers: A Survey
by: Shehzad, Ahsan, et al.
Published: (2024)
by: Shehzad, Ahsan, et al.
Published: (2024)
Evaluating AI Grading on Real-World Handwritten College Mathematics: A Large-Scale Study Toward a Benchmark
by: Yu, Zhiqi, et al.
Published: (2026)
by: Yu, Zhiqi, et al.
Published: (2026)
MH-1M: A 1.34 Million-Sample Comprehensive Multi-Feature Android Malware Dataset for Machine Learning, Deep Learning, Large Language Models, and Threat Intelligence Research
by: Braganca, Hendrio, et al.
Published: (2025)
by: Braganca, Hendrio, et al.
Published: (2025)
Authenticated Delegation and Authorized AI Agents
by: South, Tobin, et al.
Published: (2025)
by: South, Tobin, et al.
Published: (2025)
Separating Geometric Data with Minimum Cost: Two Disjoint Convex Hulls
by: Bigham, Bahram Sadeghi
Published: (2021)
by: Bigham, Bahram Sadeghi
Published: (2021)
AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization
by: Kazemi, Amir, et al.
Published: (2024)
by: Kazemi, Amir, et al.
Published: (2024)
Roughness and entropy measures of a soft set
by: Acharjee, Santanu, et al.
Published: (2026)
by: Acharjee, Santanu, et al.
Published: (2026)
Resilient Federated Chain: Transforming Blockchain Consensus into an Active Defense Layer for Federated Learning
by: García-Márquez, Mario, et al.
Published: (2026)
by: García-Márquez, Mario, et al.
Published: (2026)
Modeling Membrane Degradation in PEM Electrolyzers with Physics-Informed Neural Networks
by: Polo-Molina, Alejandro, et al.
Published: (2025)
by: Polo-Molina, Alejandro, et al.
Published: (2025)
The Hidden Costs of AI: A Review of Energy, E-Waste, and Inequality in Model Development
by: Winsta, Jenis
Published: (2025)
by: Winsta, Jenis
Published: (2025)
Temperature in SLMs: Impact on Incident Categorization in On-Premises Environments
by: Pohlmann, Marcio, et al.
Published: (2025)
by: Pohlmann, Marcio, et al.
Published: (2025)
Mathematical reasoning and the computer
by: Buzzard, Kevin
Published: (2025)
by: Buzzard, Kevin
Published: (2025)
humancompatible.detect: a Python Toolkit for Detecting Bias in AI Models
by: Matilla, German M., et al.
Published: (2025)
by: Matilla, German M., et al.
Published: (2025)
Toward a Dynamic Stackelberg Game-Theoretic Framework for Agentic AI Defense Against LLM Jailbreaking
by: Han, Zhengye, et al.
Published: (2025)
by: Han, Zhengye, et al.
Published: (2025)
Optimization before Evaluation: Evaluation with Unoptimised Prompts Can be Misleading
by: Sadjoli, Nicholas, et al.
Published: (2026)
by: Sadjoli, Nicholas, et al.
Published: (2026)
Creativity in the Age of AI: Rethinking the Role of Intentional Agency
by: Pearson, James S., et al.
Published: (2026)
by: Pearson, James S., et al.
Published: (2026)
Modeling Clinical Concern Trajectories in Language Model Agents
by: Subaharan, Sukesh, et al.
Published: (2026)
by: Subaharan, Sukesh, et al.
Published: (2026)
How well can a large language model explain business processes as perceived by users?
by: Fahland, Dirk, et al.
Published: (2024)
by: Fahland, Dirk, et al.
Published: (2024)
BernGraph: Probabilistic Graph Neural Networks for EHR-based Medication Recommendations
by: Piao, Xihao, et al.
Published: (2024)
by: Piao, Xihao, et al.
Published: (2024)
Benchmarking PNW Model for MedMNIST to 100% Accuracy
by: Deng, Bo
Published: (2026)
by: Deng, Bo
Published: (2026)
Latent-EnSF: A Latent Ensemble Score Filter for High-Dimensional Data Assimilation with Sparse Observation Data
by: Si, Phillip, et al.
Published: (2024)
by: Si, Phillip, et al.
Published: (2024)
Efficient $k$-NN Search in IoT Data: Overlap Optimization in Tree-Based Indexing Structures
by: Benrazek, Ala-Eddine, et al.
Published: (2024)
by: Benrazek, Ala-Eddine, et al.
Published: (2024)
Attention Please: What Transformer Models Really Learn for Process Prediction
by: Käppel, Martin, et al.
Published: (2024)
by: Käppel, Martin, et al.
Published: (2024)
CGRA4ML: A Hardware/Software Framework to Implement Neural Networks for Scientific Edge Computing
by: Abarajithan, G, et al.
Published: (2024)
by: Abarajithan, G, et al.
Published: (2024)
Socially Beneficial Metaverse: Framework, Technologies, Applications, and Challenges
by: Xu, Xiaolong, et al.
Published: (2023)
by: Xu, Xiaolong, et al.
Published: (2023)
A Theoretical Framework for Adaptive Utility-Weighted Benchmarking
by: Waggoner, Philip
Published: (2026)
by: Waggoner, Philip
Published: (2026)
A Human-In-The-Loop Approach for Improving Fairness in Predictive Business Process Monitoring
by: Käppel, Martin, et al.
Published: (2025)
by: Käppel, Martin, et al.
Published: (2025)
Report of the 2025 Workshop on Next-Generation Ecosystems for Scientific Computing: Harnessing Community, Software, and AI for Cross-Disciplinary Team Science
by: McInnes, Lois Curfman, et al.
Published: (2025)
by: McInnes, Lois Curfman, et al.
Published: (2025)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
by: Fang, Yunhua, et al.
Published: (2025)
by: Fang, Yunhua, et al.
Published: (2025)
Accelerating Large Kernel Convolutions with Nested Winograd Transformation.pdf
by: Jiang, Jingbo, et al.
Published: (2021)
by: Jiang, Jingbo, et al.
Published: (2021)
Isovolumetric Energy Minimization for Ball-Shaped Volume-Preserving Parameterizations of 3-Manifolds
by: Liu, Shu-Yung, et al.
Published: (2024)
by: Liu, Shu-Yung, et al.
Published: (2024)
Dynamic Observation Policies in Observation Cost-Sensitive Reinforcement Learning
by: Bellinger, Colin, et al.
Published: (2023)
by: Bellinger, Colin, et al.
Published: (2023)
Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings
by: Turan, Berkant, et al.
Published: (2025)
by: Turan, Berkant, et al.
Published: (2025)
Similar Items
-
Adaptive Orchestration for Large‐Scale Inference on Heterogeneous Accelerator Systems: Balancing Cost, Performance, and Resilience
by: Yahav Biran, et al.
Published: (2026) -
Task-parallelism in SWIFT for heterogeneous compute architectures
by: Nasar, Abouzied M. A., et al.
Published: (2025) -
Understanding Knowledge Transferability for Transfer Learning: A Survey
by: Wang, Haohua, et al.
Published: (2025) -
Deep Learning Model Deployment in Multiple Cloud Providers: an Exploratory Study Using Low Computing Power Environments
by: Lemos, Elayne, et al.
Published: (2025) -
Systems Engineering of Adaptive AI Inference Orchestration Across Heterogeneous Accelerators
by: Yahav Biran
Published: (2026)