Phantora: Maximizing Code Reuse in Simulation-based Machine Learning System Performance Estimation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qin, Jianxing, Chen, Jingrong, Kong, Xinhao, Wu, Yongji, Yuan, Tianjun, Luo, Liang, Wang, Zhaodong, Zhang, Ying, Chen, Tingjun, Lebeck, Alvin R., Zhuo, Danyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Conveyor: Efficient Tool-aware LLM Serving with Tool Partial Execution
von: Xu, Yechen, et al.
Veröffentlicht: (2024)
von: Xu, Yechen, et al.
Veröffentlicht: (2024)
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
von: Vatsavai, Sairam Sri, et al.
Veröffentlicht: (2025)
von: Vatsavai, Sairam Sri, et al.
Veröffentlicht: (2025)
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
von: McDonald, Jesse, et al.
Veröffentlicht: (2024)
von: McDonald, Jesse, et al.
Veröffentlicht: (2024)
Portable High-Performance Kernel Generation for a Computational Fluid Dynamics Code with DaCe
von: Andersson, Måns I., et al.
Veröffentlicht: (2025)
von: Andersson, Måns I., et al.
Veröffentlicht: (2025)
Temporal Load Imbalance on Ondes3D Seismic Simulator for Different Multicore Architectures
von: Solórzano, Ana Luisa Veroneze, et al.
Veröffentlicht: (2024)
von: Solórzano, Ana Luisa Veroneze, et al.
Veröffentlicht: (2024)
Shifting the Sweet Spot: High-Performance Matrix-Free Method for High-Order Elasticity
von: Chang, Dali, et al.
Veröffentlicht: (2026)
von: Chang, Dali, et al.
Veröffentlicht: (2026)
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
von: Apanasevich, L., et al.
Veröffentlicht: (2024)
von: Apanasevich, L., et al.
Veröffentlicht: (2024)
PlantD: Performance, Latency ANalysis, and Testing for Data Pipelines -- An Open Source Measurement, Testing, and Simulation Framework
von: Bogart, Christopher, et al.
Veröffentlicht: (2025)
von: Bogart, Christopher, et al.
Veröffentlicht: (2025)
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
Performance Optimization in Stream Processing Systems: Experiment-Driven Configuration Tuning for Kafka Streams
von: Chen, David, et al.
Veröffentlicht: (2026)
von: Chen, David, et al.
Veröffentlicht: (2026)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
Recorder: Comprehensive Parallel I/O Tracing and Analysis
von: Wang, Chen, et al.
Veröffentlicht: (2025)
von: Wang, Chen, et al.
Veröffentlicht: (2025)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
von: Ather, Hammad, et al.
Veröffentlicht: (2024)
von: Ather, Hammad, et al.
Veröffentlicht: (2024)
Orthrus: Accelerating Multi-BFT Consensus through Concurrent Partial Ordering of Transactions (Extended Version)
von: Lyu, Hanzheng, et al.
Veröffentlicht: (2024)
von: Lyu, Hanzheng, et al.
Veröffentlicht: (2024)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
von: Karfakis, George, et al.
Veröffentlicht: (2025)
von: Karfakis, George, et al.
Veröffentlicht: (2025)
FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
von: Xu, Jingwei, et al.
Veröffentlicht: (2025)
von: Xu, Jingwei, et al.
Veröffentlicht: (2025)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
Kairos: Efficient Temporal Graph Analytics on a Single Machine
von: da Trindade, Joana M. F., et al.
Veröffentlicht: (2024)
von: da Trindade, Joana M. F., et al.
Veröffentlicht: (2024)
Flowshop Machine Scheduling: Markov Modeling, Optimal Schedules and Heuristics
von: Ghanem, Samah A. M.
Veröffentlicht: (2025)
von: Ghanem, Samah A. M.
Veröffentlicht: (2025)
oneDAL Optimization for ARM Scalable Vector Extension: Maximizing Efficiency for High-Performance Data Science
von: Sharma, Chandan, et al.
Veröffentlicht: (2025)
von: Sharma, Chandan, et al.
Veröffentlicht: (2025)
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
von: Cornelius, Melanie, et al.
Veröffentlicht: (2025)
von: Cornelius, Melanie, et al.
Veröffentlicht: (2025)
Profiling and optimization of multi-card GPU machine learning jobs
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
Optimal Parallel Scheduling under Concave Speedup Functions
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
WebAssembly and Unikernels: A Comparative Study for Serverless at the Edge
von: Besozzi, Valerio, et al.
Veröffentlicht: (2025)
von: Besozzi, Valerio, et al.
Veröffentlicht: (2025)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes
von: Mao, Ying, et al.
Veröffentlicht: (2020)
von: Mao, Ying, et al.
Veröffentlicht: (2020)
Staging Blocked Evaluation over Structured Sparse Matrices
von: Das, Pratyush, et al.
Veröffentlicht: (2024)
von: Das, Pratyush, et al.
Veröffentlicht: (2024)
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
von: Sun, Tingyang, et al.
Veröffentlicht: (2026)
von: Sun, Tingyang, et al.
Veröffentlicht: (2026)
Reducing Tail Latencies Through Environment- and Neighbour-aware Thread Management
von: Jeffery, Andrew, et al.
Veröffentlicht: (2024)
von: Jeffery, Andrew, et al.
Veröffentlicht: (2024)
Dissecting the software-based measurement of CPU energy consumption: a comparative analysis
von: Raffin, Guillaume, et al.
Veröffentlicht: (2024)
von: Raffin, Guillaume, et al.
Veröffentlicht: (2024)
Bridding OT and PaaS in Edge-to-Cloud Continuum
von: Barrios, Carlos J, et al.
Veröffentlicht: (2025)
von: Barrios, Carlos J, et al.
Veröffentlicht: (2025)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
von: Jain, Rutwik, et al.
Veröffentlicht: (2026)
von: Jain, Rutwik, et al.
Veröffentlicht: (2026)
Hardware-Agnostic and Insightful Efficiency Metrics for Accelerated Systems: Definition and Implementation within TALP
von: Rahimi, Ghazal, et al.
Veröffentlicht: (2026)
von: Rahimi, Ghazal, et al.
Veröffentlicht: (2026)
Node Compass: Multilevel Tracing and Debugging of Request Executions in JavaScript-Based Web-Servers
von: Kabamba, Herve Mbikayi, et al.
Veröffentlicht: (2023)
von: Kabamba, Herve Mbikayi, et al.
Veröffentlicht: (2023)
Beyond Thread States: Diagnosing Performance Degradation with eBPF and Thread Dynamics
von: Landau, Diogo, et al.
Veröffentlicht: (2026)
von: Landau, Diogo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Conveyor: Efficient Tool-aware LLM Serving with Tool Partial Execution
von: Xu, Yechen, et al.
Veröffentlicht: (2024) -
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
von: Vatsavai, Sairam Sri, et al.
Veröffentlicht: (2025) -
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
von: McDonald, Jesse, et al.
Veröffentlicht: (2024) -
Portable High-Performance Kernel Generation for a Computational Fluid Dynamics Code with DaCe
von: Andersson, Måns I., et al.
Veröffentlicht: (2025) -
Temporal Load Imbalance on Ondes3D Seismic Simulator for Different Multicore Architectures
von: Solórzano, Ana Luisa Veroneze, et al.
Veröffentlicht: (2024)