Comprehensive Performance Modeling and System Design Insights for Foundation Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Subramanian, Shashank, Rrapaj, Ermal, Harrington, Peter, Chheda, Smeet, Farrell, Steven, Austin, Brian, Williams, Samuel, Wright, Nicholas, Bhimji, Wahid |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
von: Golden, Alicia, et al.
Veröffentlicht: (2025)
von: Golden, Alicia, et al.
Veröffentlicht: (2025)
EvoSort: A Genetic-Algorithm-Based Adaptive Parallel Sorting Framework for Large-Scale High Performance Computing
von: Raj, Shashank, et al.
Veröffentlicht: (2025)
von: Raj, Shashank, et al.
Veröffentlicht: (2025)
The Time to Consensus in a Blockchain: Insights into Bitcoin's "6 Blocks Rule''
von: Dey, Partha S., et al.
Veröffentlicht: (2025)
von: Dey, Partha S., et al.
Veröffentlicht: (2025)
A New Execution Model and Executor for Adaptively Optimizing the Performance of Parallel Algorithms Using HPX Runtime System
von: Mohammadiporshokooh, Karame, et al.
Veröffentlicht: (2025)
von: Mohammadiporshokooh, Karame, et al.
Veröffentlicht: (2025)
Characterizing Production GPU Workloads using System-wide Telemetry Data
von: Cankur, Onur, et al.
Veröffentlicht: (2025)
von: Cankur, Onur, et al.
Veröffentlicht: (2025)
Beyond Pre-Training: The Full Lifecycle of Foundation Models on HPC Systems
von: Conciatore, Dino, et al.
Veröffentlicht: (2026)
von: Conciatore, Dino, et al.
Veröffentlicht: (2026)
Modular Foundation Model Inference at the Edge: Network-Aware Microservice Optimization
von: Zhu, Juan, et al.
Veröffentlicht: (2026)
von: Zhu, Juan, et al.
Veröffentlicht: (2026)
LAPIS: A Performance Portable, High Productivity Compiler Framework
von: Kelley, Brian, et al.
Veröffentlicht: (2025)
von: Kelley, Brian, et al.
Veröffentlicht: (2025)
Unleashing Scalable Context Parallelism for Foundation Models Pre-Training via FCP
von: Zhao, Yilong, et al.
Veröffentlicht: (2026)
von: Zhao, Yilong, et al.
Veröffentlicht: (2026)
Large Scale Multi-GPU Based Parallel Traffic Simulation for Accelerated Traffic Assignment and Propagation
von: Jiang, Xuan, et al.
Veröffentlicht: (2024)
von: Jiang, Xuan, et al.
Veröffentlicht: (2024)
A Comprehensive Hyperledger Fabric Performance Evaluation based on Resources Capacity Planning
von: Melo, Carlos, et al.
Veröffentlicht: (2025)
von: Melo, Carlos, et al.
Veröffentlicht: (2025)
Benchmarking Message Brokers for IoT Edge Computing: A Comprehensive Performance Study
von: Paul, Tapajit Chandra, et al.
Veröffentlicht: (2026)
von: Paul, Tapajit Chandra, et al.
Veröffentlicht: (2026)
Rorqual: Speeding up Narwhal with TEEs
von: Freitas, Luciano, et al.
Veröffentlicht: (2024)
von: Freitas, Luciano, et al.
Veröffentlicht: (2024)
System-Level Performance Modeling of Photonic In-Memory Computing
von: Arockiaraj, Jebacyril, et al.
Veröffentlicht: (2026)
von: Arockiaraj, Jebacyril, et al.
Veröffentlicht: (2026)
Performance Models for a Two-tiered Storage System
von: Sasidharan, Aparna, et al.
Veröffentlicht: (2025)
von: Sasidharan, Aparna, et al.
Veröffentlicht: (2025)
ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace
von: Shi, Ruimin, et al.
Veröffentlicht: (2025)
von: Shi, Ruimin, et al.
Veröffentlicht: (2025)
Optimal Resource Utilization in Hyperledger Fabric: A Comprehensive SPN-Based Performance Evaluation Paradigm
von: Melo, Carlos, et al.
Veröffentlicht: (2025)
von: Melo, Carlos, et al.
Veröffentlicht: (2025)
Training DNN Models over Heterogeneous Clusters with Optimal Performance
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
Empowering the Quantum Cloud User with QRIO
von: Chakraborty, Shmeelok, et al.
Veröffentlicht: (2024)
von: Chakraborty, Shmeelok, et al.
Veröffentlicht: (2024)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
von: Kumar, Madabattula Rajesh, et al.
Veröffentlicht: (2025)
von: Kumar, Madabattula Rajesh, et al.
Veröffentlicht: (2025)
Evaluation of Programming Models and Performance for Stencil Computation on Current GPU Architectures
von: Shan, Baodi, et al.
Veröffentlicht: (2024)
von: Shan, Baodi, et al.
Veröffentlicht: (2024)
Benchmarking the Performance of Large Language Models on the Cerebras Wafer Scale Engine
von: Zhang, Zuoning, et al.
Veröffentlicht: (2024)
von: Zhang, Zuoning, et al.
Veröffentlicht: (2024)
Experiences with Model Context Protocol Servers for Science and High Performance Computing
von: Pan, Haochen, et al.
Veröffentlicht: (2025)
von: Pan, Haochen, et al.
Veröffentlicht: (2025)
MoFa: A Unified Performance Modeling Framework for LLM Pretraining
von: Zhao, Lu, et al.
Veröffentlicht: (2025)
von: Zhao, Lu, et al.
Veröffentlicht: (2025)
Privacy-Preserving Sharing of Data Analytics Runtime Metrics for Performance Modeling
von: Will, Jonathan, et al.
Veröffentlicht: (2024)
von: Will, Jonathan, et al.
Veröffentlicht: (2024)
PICO: Performance Insights for Collective Operations
von: Pasqualoni, Saverio, et al.
Veröffentlicht: (2025)
von: Pasqualoni, Saverio, et al.
Veröffentlicht: (2025)
Design Principles of Dynamic Resource Management for High-Performance Parallel Programming Models
von: Huber, Dominik, et al.
Veröffentlicht: (2024)
von: Huber, Dominik, et al.
Veröffentlicht: (2024)
Characterizing the Performance of Accelerated Jetson Edge Devices for Training Deep Learning Models
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
Efficient Training Approaches for Performance Anomaly Detection Models in Edge Computing Environments
von: Fernando, Duneesha, et al.
Veröffentlicht: (2024)
von: Fernando, Duneesha, et al.
Veröffentlicht: (2024)
ML-based Modeling to Predict I/O Performance on Different Storage Sub-systems
von: Xu, Yiheng, et al.
Veröffentlicht: (2023)
von: Xu, Yiheng, et al.
Veröffentlicht: (2023)
Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO
von: Svedas, Jonas, et al.
Veröffentlicht: (2026)
von: Svedas, Jonas, et al.
Veröffentlicht: (2026)
Parallel Reduced Order Modeling for Digital Twins using High-Performance Computing Workflows
von: de Parga, S. Ares, et al.
Veröffentlicht: (2024)
von: de Parga, S. Ares, et al.
Veröffentlicht: (2024)
Performance Modeling and Evaluation of Hyperledger Fabric: An Analysis Based on Transaction Flow and Endorsement Policies
von: Melo, Carlos, et al.
Veröffentlicht: (2025)
von: Melo, Carlos, et al.
Veröffentlicht: (2025)
PALM: A Efficient Performance Simulator for Tiled Accelerators with Large-scale Model Training
von: Fang, Jiahao, et al.
Veröffentlicht: (2024)
von: Fang, Jiahao, et al.
Veröffentlicht: (2024)
Transactional Dynamics in Hyperledger Fabric: A Stochastic Modeling and Performance Evaluation of Permissioned Blockchains
von: Melo, Carlos, et al.
Veröffentlicht: (2025)
von: Melo, Carlos, et al.
Veröffentlicht: (2025)
Cascadia: An Efficient Cascade Serving System for Large Language Models
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
Leveraging HPC Profiling & Tracing Tools to Understand the Performance of Particle-in-Cell Monte Carlo Simulations
von: Williams, Jeremy J., et al.
Veröffentlicht: (2023)
von: Williams, Jeremy J., et al.
Veröffentlicht: (2023)
M3SA: Exploring Datacenter Performance and Climate-Impact with Multi- and Meta-Model Simulation and Analysis
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
LoHan: Low-Cost High-Performance Framework to Fine-Tune 100B Model on a Consumer GPU
von: Liao, Changyue, et al.
Veröffentlicht: (2024)
von: Liao, Changyue, et al.
Veröffentlicht: (2024)
Optimizing Data Distribution and Kernel Performance for Efficient Training of Chemistry Foundation Models: A Case Study with MACE
von: Firoz, Jesun, et al.
Veröffentlicht: (2025)
von: Firoz, Jesun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
von: Golden, Alicia, et al.
Veröffentlicht: (2025) -
EvoSort: A Genetic-Algorithm-Based Adaptive Parallel Sorting Framework for Large-Scale High Performance Computing
von: Raj, Shashank, et al.
Veröffentlicht: (2025) -
The Time to Consensus in a Blockchain: Insights into Bitcoin's "6 Blocks Rule''
von: Dey, Partha S., et al.
Veröffentlicht: (2025) -
A New Execution Model and Executor for Adaptively Optimizing the Performance of Parallel Algorithms Using HPX Runtime System
von: Mohammadiporshokooh, Karame, et al.
Veröffentlicht: (2025) -
Characterizing Production GPU Workloads using System-wide Telemetry Data
von: Cankur, Onur, et al.
Veröffentlicht: (2025)