Experience Deploying Containerized GenAI Services at an HPC Center
Fuente:
arXiv
Saved in:
| Main Authors: | Beltre, Angel M., Ogden, Jeff, Pedretti, Kevin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Intelligence: Designing Data Centers for Next-Gen Language Models
by: Tithi, Jesmin Jahan, et al.
Published: (2025)
by: Tithi, Jesmin Jahan, et al.
Published: (2025)
Managed-Retention Memory: A New Class of Memory for the AI Era
by: Legtchenko, Sergey, et al.
Published: (2025)
by: Legtchenko, Sergey, et al.
Published: (2025)
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference
by: Ortega, Cristobal, et al.
Published: (2024)
by: Ortega, Cristobal, et al.
Published: (2024)
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
by: Afzal, Ayesha, et al.
Published: (2026)
by: Afzal, Ayesha, et al.
Published: (2026)
CLAASIC: a Cortex-Inspired Hardware Accelerator
by: Puente, Valentin, et al.
Published: (2016)
by: Puente, Valentin, et al.
Published: (2016)
WaferLLM: Large Language Model Inference at Wafer Scale
by: He, Congjie, et al.
Published: (2025)
by: He, Congjie, et al.
Published: (2025)
Transforming the Hybrid Cloud for Emerging AI Workloads
by: Chen, Deming, et al.
Published: (2024)
by: Chen, Deming, et al.
Published: (2024)
Open Challenges for a Production-ready Cloud Environment on top of RISC-V hardware
by: Call, Aaron, et al.
Published: (2025)
by: Call, Aaron, et al.
Published: (2025)
From GPUs to RRAMs: Distributed In-Memory Primal-Dual Hybrid Gradient Method for Solving Large-Scale Linear Optimization Problem
by: Vo, Huynh Q. N., et al.
Published: (2025)
by: Vo, Huynh Q. N., et al.
Published: (2025)
DFabric: Scaling Out Data Parallel Applications with CXL-Ethernet Hybrid Interconnects
by: Zhang, Xu, et al.
Published: (2024)
by: Zhang, Xu, et al.
Published: (2024)
Carbon Connect: An Ecosystem for Sustainable Computing
by: Lee, Benjamin C., et al.
Published: (2024)
by: Lee, Benjamin C., et al.
Published: (2024)
Efficient Optimization Accelerator Framework for Multistate Ising Problems
by: Garg, Chirag, et al.
Published: (2025)
by: Garg, Chirag, et al.
Published: (2025)
TreeVQA: A Tree-Structured Execution Framework for Shot Reduction in Variational Quantum Algorithms
by: Hou, Yuewen, et al.
Published: (2025)
by: Hou, Yuewen, et al.
Published: (2025)
Architecting Distributed Quantum Computers: Design Insights from Resource Estimation
by: Filippov, Dmitry, et al.
Published: (2025)
by: Filippov, Dmitry, et al.
Published: (2025)
ForgetMeNot: Understanding and Modeling the Impact of Forever Chemicals Toward Sustainable Large-Scale Computing
by: Roy, Rohan Basu, et al.
Published: (2025)
by: Roy, Rohan Basu, et al.
Published: (2025)
Reference Architecture of a Quantum-Centric Supercomputer
by: Seelam, Seetharami, et al.
Published: (2026)
by: Seelam, Seetharami, et al.
Published: (2026)
Sustainable Supercomputing for AI: GPU Power Capping at HPC Scale
by: Zhao, Dan, et al.
Published: (2024)
by: Zhao, Dan, et al.
Published: (2024)
COMPASS: A Compiler Framework for Resource-Constrained Crossbar-Array Based In-Memory Deep Learning Accelerators
by: Park, Jihoon, et al.
Published: (2025)
by: Park, Jihoon, et al.
Published: (2025)
Evaluating Kubernetes Performance for GenAI Inference: From Automatic Speech Recognition to LLM Summarization
by: Malleni, Sai Sindhur, et al.
Published: (2026)
by: Malleni, Sai Sindhur, et al.
Published: (2026)
Deep Tech to Space: Space Data Centers and AI Revolution at the Edge
by: Weiss, Jonas, et al.
Published: (2026)
by: Weiss, Jonas, et al.
Published: (2026)
Harnessing the Full Potential of RRAMs through Scalable and Distributed In-Memory Computing with Integrated Error Correction
by: Vo, Huynh Q. N., et al.
Published: (2025)
by: Vo, Huynh Q. N., et al.
Published: (2025)
Efficient Edge AI: Deploying Convolutional Neural Networks on FPGA with the Gemmini Accelerator
by: Peccia, Federico Nicolas, et al.
Published: (2024)
by: Peccia, Federico Nicolas, et al.
Published: (2024)
Application Experiences on a GPU-Accelerated Arm-based HPC Testbed
by: Elwasif, Wael, et al.
Published: (2022)
by: Elwasif, Wael, et al.
Published: (2022)
EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
by: Kubwimana, Benjamin, et al.
Published: (2025)
by: Kubwimana, Benjamin, et al.
Published: (2025)
PASS: An Asynchronous Probabilistic Processor for Next Generation Intelligence
by: Patel, Saavan, et al.
Published: (2024)
by: Patel, Saavan, et al.
Published: (2024)
Profiling AI Models: Towards Efficient Computation Offloading in Heterogeneous Edge AI Systems
by: Parra-Ullauri, Juan Marcelo, et al.
Published: (2024)
by: Parra-Ullauri, Juan Marcelo, et al.
Published: (2024)
Power Stabilization for AI Training Datacenters
by: Choukse, Esha, et al.
Published: (2025)
by: Choukse, Esha, et al.
Published: (2025)
GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
by: Kumar, Deepak, et al.
Published: (2025)
by: Kumar, Deepak, et al.
Published: (2025)
Heterogeneous Computing: The Key to Powering the Future of AI Agent Inference
by: Zhao, Yiren, et al.
Published: (2026)
by: Zhao, Yiren, et al.
Published: (2026)
Debunking the CUDA Myth Towards GPU-based AI Systems
by: Lee, Yunjae, et al.
Published: (2024)
by: Lee, Yunjae, et al.
Published: (2024)
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
by: Stojkovic, Jovan, et al.
Published: (2025)
by: Stojkovic, Jovan, et al.
Published: (2025)
Improving AI Efficiency in Data Centres by Power Dynamic Response
by: Marinoni, Andrea, et al.
Published: (2025)
by: Marinoni, Andrea, et al.
Published: (2025)
Modernizing Amdahl's Law: How AI Scaling Laws Shape Computer Architecture
by: Lu, Chien-Ping
Published: (2026)
by: Lu, Chien-Ping
Published: (2026)
ZettaLith: An Architectural Exploration of Extreme-Scale AI Inference Acceleration
by: Silverbrook, Kia
Published: (2025)
by: Silverbrook, Kia
Published: (2025)
Exploring energy consumption of AI frameworks on a 64-core RV64 Server CPU
by: Malenza, Giulio, et al.
Published: (2025)
by: Malenza, Giulio, et al.
Published: (2025)
Good things come in small packages: Should we build AI clusters with Lite-GPUs?
by: Canakci, Burcu, et al.
Published: (2025)
by: Canakci, Burcu, et al.
Published: (2025)
The DMA Streaming Framework: Kernel-Level Buffer Orchestration for High-Performance AI Data Paths
by: Graziano, Marco
Published: (2026)
by: Graziano, Marco
Published: (2026)
Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models
by: Bambhaniya, Abhimanyu, et al.
Published: (2024)
by: Bambhaniya, Abhimanyu, et al.
Published: (2024)
Advancing AI-assisted Hardware Design with Hierarchical Decentralized Training and Personalized Inference-Time Optimization
by: Chen, Hao Mark, et al.
Published: (2025)
by: Chen, Hao Mark, et al.
Published: (2025)
Sustainable AI Training via Hardware-Software Co-Design on NVIDIA, AMD, and Emerging GPU Architectures
by: Makin, Yashasvi, et al.
Published: (2025)
by: Makin, Yashasvi, et al.
Published: (2025)
Similar Items
-
Scaling Intelligence: Designing Data Centers for Next-Gen Language Models
by: Tithi, Jesmin Jahan, et al.
Published: (2025) -
Managed-Retention Memory: A New Class of Memory for the AI Era
by: Legtchenko, Sergey, et al.
Published: (2025) -
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference
by: Ortega, Cristobal, et al.
Published: (2024) -
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
by: Afzal, Ayesha, et al.
Published: (2026) -
CLAASIC: a Cortex-Inspired Hardware Accelerator
by: Puente, Valentin, et al.
Published: (2016)