Scalable Engine and the Performance of Different LLM Models in a SLURM based HPC architecture
Fuente:
arXiv
Salvato in:
| Autori principali: | Luiz, Anderson de Lima, Kurlekar, Shubham Vijay, Georges, Munir |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs
di: Sada, Mohammad Firas, et al.
Pubblicazione: (2025)
di: Sada, Mohammad Firas, et al.
Pubblicazione: (2025)
Efficiently Scheduling Parallel DAG Tasks on Identical Multiprocessors
di: Lendve, Shardul, et al.
Pubblicazione: (2024)
di: Lendve, Shardul, et al.
Pubblicazione: (2024)
A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models
di: Zhang, Lingzhe, et al.
Pubblicazione: (2025)
di: Zhang, Lingzhe, et al.
Pubblicazione: (2025)
Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
di: Li, Zonghang, et al.
Pubblicazione: (2025)
di: Li, Zonghang, et al.
Pubblicazione: (2025)
On-Device Generative AI for GDPR-Compliant Visual Monitoring: Natural Language Alerts from Local Object Detection
di: Schappacher-Tilp, Gudrun, et al.
Pubblicazione: (2026)
di: Schappacher-Tilp, Gudrun, et al.
Pubblicazione: (2026)
DPDPU: Data Processing with DPUs
di: Hu, Jiasheng, et al.
Pubblicazione: (2024)
di: Hu, Jiasheng, et al.
Pubblicazione: (2024)
Flex-MIG: Enabling Distributed Execution on MIG
di: Kim, Myeongsu, et al.
Pubblicazione: (2025)
di: Kim, Myeongsu, et al.
Pubblicazione: (2025)
ITQ3_S: High-Fidelity 3-bit LLM Inference via Interleaved Ternary Quantization with Rotation-Domain Smoothing
di: Yoon, Edward J.
Pubblicazione: (2026)
di: Yoon, Edward J.
Pubblicazione: (2026)
push0: Scalable and Fault-Tolerant Orchestration for Zero-Knowledge Proof Generation
di: Ahmadvand, Mohsen, et al.
Pubblicazione: (2026)
di: Ahmadvand, Mohsen, et al.
Pubblicazione: (2026)
CRDT-Based Game State Synchronization in Peer-to-Peer VR
di: Dantas, Abel, et al.
Pubblicazione: (2025)
di: Dantas, Abel, et al.
Pubblicazione: (2025)
GraphBit: A Graph-based Agentic Framework for Non-Linear Agent Orchestration
di: Sarker, Yeahia, et al.
Pubblicazione: (2026)
di: Sarker, Yeahia, et al.
Pubblicazione: (2026)
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
di: Li, Xiangchen, et al.
Pubblicazione: (2026)
di: Li, Xiangchen, et al.
Pubblicazione: (2026)
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
di: Li, Xiangchen, et al.
Pubblicazione: (2026)
di: Li, Xiangchen, et al.
Pubblicazione: (2026)
nvidia-pcm: A D-Bus-Driven Platform Configuration Manager for OpenBMC Environments
di: Singh, Harinder
Pubblicazione: (2026)
di: Singh, Harinder
Pubblicazione: (2026)
Serverless Cold Starts and Where to Find Them
di: Joosen, Artjom, et al.
Pubblicazione: (2024)
di: Joosen, Artjom, et al.
Pubblicazione: (2024)
ZenFlow: Enabling Stall-Free Offloading Training via Asynchronous Updates
di: Lan, Tingfeng, et al.
Pubblicazione: (2025)
di: Lan, Tingfeng, et al.
Pubblicazione: (2025)
D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree
di: Lei, Xiang, et al.
Pubblicazione: (2025)
di: Lei, Xiang, et al.
Pubblicazione: (2025)
Melding the Serverless Control Plane with the Conventional Cluster Manager for Speed and Resource Efficiency
di: Kondrashov, Leonid, et al.
Pubblicazione: (2025)
di: Kondrashov, Leonid, et al.
Pubblicazione: (2025)
Reexamining Paradigms of End-to-End Data Movement
di: Fang, Chin, et al.
Pubblicazione: (2025)
di: Fang, Chin, et al.
Pubblicazione: (2025)
Addressing tokens dynamic generation, propagation, storage and renewal to secure the GlideinWMS pilot based jobs and system
di: Coimbra, Bruno Moreira, et al.
Pubblicazione: (2025)
di: Coimbra, Bruno Moreira, et al.
Pubblicazione: (2025)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
di: Kolluru, Saicharan
Pubblicazione: (2025)
di: Kolluru, Saicharan
Pubblicazione: (2025)
Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning
di: Zhang, Li, et al.
Pubblicazione: (2025)
di: Zhang, Li, et al.
Pubblicazione: (2025)
Sure! Here's a short and concise title for your paper: "Contamination in Generated Text Detection Benchmarks"
di: Dingfelder, Philipp, et al.
Pubblicazione: (2025)
di: Dingfelder, Philipp, et al.
Pubblicazione: (2025)
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
di: Agarwal, Amit, et al.
Pubblicazione: (2025)
di: Agarwal, Amit, et al.
Pubblicazione: (2025)
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
di: Kamath, Aditya K, et al.
Pubblicazione: (2024)
di: Kamath, Aditya K, et al.
Pubblicazione: (2024)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
di: Penke, Carolin, et al.
Pubblicazione: (2025)
di: Penke, Carolin, et al.
Pubblicazione: (2025)
Advancing Annotat3D with Harpia: A CUDA-Accelerated Library For Large-Scale Volumetric Data Segmentation
di: de Araujo, Camila Machado, et al.
Pubblicazione: (2025)
di: de Araujo, Camila Machado, et al.
Pubblicazione: (2025)
Deep RC: A Scalable Data Engineering and Deep Learning Pipeline
di: Sarker, Arup Kumar, et al.
Pubblicazione: (2025)
di: Sarker, Arup Kumar, et al.
Pubblicazione: (2025)
Evaluating Large Language Models for Workload Mapping and Scheduling in Heterogeneous HPC Systems
di: Sharma, Aasish Kumar, et al.
Pubblicazione: (2025)
di: Sharma, Aasish Kumar, et al.
Pubblicazione: (2025)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
di: Haque, Md. Asraful, et al.
Pubblicazione: (2026)
di: Haque, Md. Asraful, et al.
Pubblicazione: (2026)
The High Cost of Keeping Warm: Characterizing Overhead in Serverless Autoscaling Policies
di: Kondrashov, Leonid, et al.
Pubblicazione: (2025)
di: Kondrashov, Leonid, et al.
Pubblicazione: (2025)
Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data
di: Borisov, Vadim
Pubblicazione: (2026)
di: Borisov, Vadim
Pubblicazione: (2026)
Retrieval-Based Multi-Label Legal Annotation: Extensible, Data-Efficient and Hallucination-Free
di: Zhang, Li, et al.
Pubblicazione: (2026)
di: Zhang, Li, et al.
Pubblicazione: (2026)
SAGED: A Holistic Bias-Benchmarking Pipeline for Language Models with Customisable Fairness Calibration
di: Guan, Xin, et al.
Pubblicazione: (2024)
di: Guan, Xin, et al.
Pubblicazione: (2024)
Accelerating In-transit Isosurface Generation With Topology Preserving Compression
di: Li, Yanliang, et al.
Pubblicazione: (2024)
di: Li, Yanliang, et al.
Pubblicazione: (2024)
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended
di: Kamath, Aditya K, et al.
Pubblicazione: (2026)
di: Kamath, Aditya K, et al.
Pubblicazione: (2026)
Shattering the Ephemeral Storage Cost Barrier for Data-Intensive Serverless Workflows
di: Ustiugov, Dmitrii, et al.
Pubblicazione: (2023)
di: Ustiugov, Dmitrii, et al.
Pubblicazione: (2023)
Towards Robust Retrieval-Augmented Generation Based on Knowledge Graph: A Comparative Analysis
di: Amamou, Hazem, et al.
Pubblicazione: (2026)
di: Amamou, Hazem, et al.
Pubblicazione: (2026)
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
di: Georgiou, Athos
Pubblicazione: (2026)
di: Georgiou, Athos
Pubblicazione: (2026)
ReliabilityBench: Evaluating LLM Agent Reliability Under Production-Like Stress Conditions
di: Gupta, Aayush
Pubblicazione: (2026)
di: Gupta, Aayush
Pubblicazione: (2026)
Documenti analoghi
-
Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs
di: Sada, Mohammad Firas, et al.
Pubblicazione: (2025) -
Efficiently Scheduling Parallel DAG Tasks on Identical Multiprocessors
di: Lendve, Shardul, et al.
Pubblicazione: (2024) -
A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models
di: Zhang, Lingzhe, et al.
Pubblicazione: (2025) -
Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
di: Li, Zonghang, et al.
Pubblicazione: (2025) -
On-Device Generative AI for GDPR-Compliant Visual Monitoring: Natural Language Alerts from Local Object Detection
di: Schappacher-Tilp, Gudrun, et al.
Pubblicazione: (2026)