Chat AI: A Seamless Slurm-Native Solution for HPC-Based Services
Fuente:
arXiv
Saved in:
| Main Authors: | Doosthosseini, Ali, Decker, Jonathan, Nolte, Hendrik, Kunkel, Julian M. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Review of Tools and Techniques for Optimization of Workload Mapping and Scheduling in Heterogeneous HPC System
by: Sharma, Aasish Kumar, et al.
Published: (2025)
by: Sharma, Aasish Kumar, et al.
Published: (2025)
DECICE: AI-Driven Scheduling and Digital Twin Integration for the Cloud-HPC-Edge Compute Continuum
by: Sharma, Aasish Kumar, et al.
Published: (2026)
by: Sharma, Aasish Kumar, et al.
Published: (2026)
AI Factories: It's time to rethink the Cloud-HPC divide
by: Lopez, Pedro Garcia, et al.
Published: (2025)
by: Lopez, Pedro Garcia, et al.
Published: (2025)
LLM as HPC Expert: Extending RAG Architecture for HPC Data
by: Miyashita, Yusuke, et al.
Published: (2024)
by: Miyashita, Yusuke, et al.
Published: (2024)
Evaluating the Efficacy of LLM-Based Reasoning for Multiobjective HPC Job Scheduling
by: Jadhav, Prachi, et al.
Published: (2025)
by: Jadhav, Prachi, et al.
Published: (2025)
Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep Learning
by: An, Wei, et al.
Published: (2024)
by: An, Wei, et al.
Published: (2024)
Survey of HPC in US Research Institutions
by: Shu, Peng, et al.
Published: (2025)
by: Shu, Peng, et al.
Published: (2025)
Automated Dynamic AI Inference Scaling on HPC-Infrastructure: Integrating Kubernetes, Slurm and vLLM
by: Trappen, Tim, et al.
Published: (2025)
by: Trappen, Tim, et al.
Published: (2025)
GrapheonRL: A Graph Neural Network and Reinforcement Learning Framework for Constraint and Data-Aware Workflow Mapping and Scheduling in Heterogeneous HPC Systems
by: Sharma, Aasish Kumar, et al.
Published: (2025)
by: Sharma, Aasish Kumar, et al.
Published: (2025)
Block size estimation for data partitioning in HPC applications using machine learning techniques
by: Cantini, Riccardo, et al.
Published: (2022)
by: Cantini, Riccardo, et al.
Published: (2022)
HPC-Coder: Modeling Parallel Programs using Large Language Models
by: Nichols, Daniel, et al.
Published: (2023)
by: Nichols, Daniel, et al.
Published: (2023)
Enabling Seamless Transitions from Experimental to Production HPC for Interactive Workflows
by: Etz, Brian D., et al.
Published: (2025)
by: Etz, Brian D., et al.
Published: (2025)
LLM & HPC:Benchmarking DeepSeek's Performance in High-Performance Computing Tasks
by: Nader, Noujoud, et al.
Published: (2025)
by: Nader, Noujoud, et al.
Published: (2025)
Scalable Runtime Architecture for Data-driven, Hybrid HPC and ML Workflow Applications
by: Merzky, Andre, et al.
Published: (2025)
by: Merzky, Andre, et al.
Published: (2025)
ClusterRCA: An End-to-End Approach for Network Fault Localization and Classification for HPC System
by: Sun, Yongqian, et al.
Published: (2025)
by: Sun, Yongqian, et al.
Published: (2025)
Cloud Platforms for Developing Generative AI Solutions: A Scoping Review of Tools and Services
by: Patel, Dhavalkumar, et al.
Published: (2024)
by: Patel, Dhavalkumar, et al.
Published: (2024)
Sustainable Supercomputing for AI: GPU Power Capping at HPC Scale
by: Zhao, Dan, et al.
Published: (2024)
by: Zhao, Dan, et al.
Published: (2024)
Harnessing Deep Learning and HPC Kernels via High-Level Loop and Tensor Abstractions on CPU Architectures
by: Georganas, Evangelos, et al.
Published: (2023)
by: Georganas, Evangelos, et al.
Published: (2023)
VibeCodeHPC: An Agent-Based Iterative Prompting Auto-Tuner for HPC Code Generation Using LLMs
by: Hayashi, Shun-ichiro, et al.
Published: (2025)
by: Hayashi, Shun-ichiro, et al.
Published: (2025)
Analytically-Driven Resource Management for Cloud-Native Microservices
by: Zhang, Yanqi, et al.
Published: (2024)
by: Zhang, Yanqi, et al.
Published: (2024)
Towards an Adaptive Runtime System for Cloud-Native HPC
by: Bhosale, Aditya, et al.
Published: (2026)
by: Bhosale, Aditya, et al.
Published: (2026)
LlamaDuo: LLMOps Pipeline for Seamless Migration from Service LLMs to Small-Scale Local LLMs
by: Park, Chansung, et al.
Published: (2024)
by: Park, Chansung, et al.
Published: (2024)
Scalable Cloud-Native Architectures for Intelligent PMU Data Processing
by: Chockalingam, Nachiappan, et al.
Published: (2025)
by: Chockalingam, Nachiappan, et al.
Published: (2025)
LLMs as Packagers of HPC Software
by: Melone, Caetano, et al.
Published: (2025)
by: Melone, Caetano, et al.
Published: (2025)
Artifact for Service-Level Energy Modeling and Experimentation for Cloud-Native Microservices
by: Legler, Julian
Published: (2026)
by: Legler, Julian
Published: (2026)
Thousand-GPU Large-Scale Training and Optimization Recipe for AI-Native Cloud Embodied Intelligence Infrastructure
by: Guo, Yongjian, et al.
Published: (2026)
by: Guo, Yongjian, et al.
Published: (2026)
Ontological Knowledge Blocks: Executable Compliance and Profile-Based Validation for Trustworthy AI Systems
by: Sharma, Aasish Kumar, et al.
Published: (2026)
by: Sharma, Aasish Kumar, et al.
Published: (2026)
Experience Deploying Containerized GenAI Services at an HPC Center
by: Beltre, Angel M., et al.
Published: (2025)
by: Beltre, Angel M., et al.
Published: (2025)
Leveraging AI for Productive and Trustworthy HPC Software: Challenges and Research Directions
by: Teranishi, Keita, et al.
Published: (2025)
by: Teranishi, Keita, et al.
Published: (2025)
Elastic On-Device LLM Service
by: Yin, Wangsong, et al.
Published: (2024)
by: Yin, Wangsong, et al.
Published: (2024)
Reconstruction-Based Adaptive Scheduling Using AI Inferences in Safety-Critical Systems
by: Alshaer, Samer, et al.
Published: (2025)
by: Alshaer, Samer, et al.
Published: (2025)
Deploying Foundation Model Powered Agent Services: A Survey
by: Xu, Wenchao, et al.
Published: (2024)
by: Xu, Wenchao, et al.
Published: (2024)
Solutions for Distributed Memory Access Mechanism on HPC Clusters
by: Meizner, Jan, et al.
Published: (2025)
by: Meizner, Jan, et al.
Published: (2025)
Integrating and Characterizing HPC Task Runtime Systems for hybrid AI-HPC workloads
by: Merzky, Andre, et al.
Published: (2025)
by: Merzky, Andre, et al.
Published: (2025)
Service-Level Energy Modeling and Experimentation for Cloud-Native Microservices
by: Legler, Julian, et al.
Published: (2025)
by: Legler, Julian, et al.
Published: (2025)
Cloud-Based AI Systems: Leveraging Large Language Models for Intelligent Fault Detection and Autonomous Self-Healing
by: Ji, Cheng, et al.
Published: (2025)
by: Ji, Cheng, et al.
Published: (2025)
Enhancing Cluster Scheduling in HPC: A Continuous Transfer Learning for Real-Time Optimization
by: Sliwko, Leszek, et al.
Published: (2025)
by: Sliwko, Leszek, et al.
Published: (2025)
From Edge to HPC: Investigating Cross-Facility Data Streaming Architectures
by: George, Anjus, et al.
Published: (2025)
by: George, Anjus, et al.
Published: (2025)
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
by: Yu, Dianhai, et al.
Published: (2022)
by: Yu, Dianhai, et al.
Published: (2022)
MARLaaS: Multi-Tenant Asynchronous Reinforcement Learning as a Service
by: Yu, Timothy Tin Long, et al.
Published: (2026)
by: Yu, Timothy Tin Long, et al.
Published: (2026)
Similar Items
-
A Review of Tools and Techniques for Optimization of Workload Mapping and Scheduling in Heterogeneous HPC System
by: Sharma, Aasish Kumar, et al.
Published: (2025) -
DECICE: AI-Driven Scheduling and Digital Twin Integration for the Cloud-HPC-Edge Compute Continuum
by: Sharma, Aasish Kumar, et al.
Published: (2026) -
AI Factories: It's time to rethink the Cloud-HPC divide
by: Lopez, Pedro Garcia, et al.
Published: (2025) -
LLM as HPC Expert: Extending RAG Architecture for HPC Data
by: Miyashita, Yusuke, et al.
Published: (2024) -
Evaluating the Efficacy of LLM-Based Reasoning for Multiobjective HPC Job Scheduling
by: Jadhav, Prachi, et al.
Published: (2025)