MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sridharan, Srinivas, Badea, Theodor-Adrian, Balogh, Andy, Beckmann, Bradford M., Coutinho, Brian, Feng, Louis, Fu, Sheng, Gao, Sanshan, Garakani, Mehryar, Heo, Taekyung, Kanter, David, Ladd, Josh, Li, Ziwei, Liu, Winston, Man, Changhai, Mihailescu, Dan, More, Spandan, Park, Joongun, Ramachandran, Ashwin, Ramakrishnaiah, Vinay, Rashidi, Saeed, Reddi, Vijay Janapa, Sharma, Puneet, Tian, Phio, Won, William, Wu, Hanjiang, Xu, Huan, Yoo, Jinsun, Krishna, Tushar
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913143251271680
author Sridharan, Srinivas
Badea, Theodor-Adrian
Balogh, Andy
Beckmann, Bradford M.
Coutinho, Brian
Feng, Louis
Fu, Sheng
Gao, Sanshan
Garakani, Mehryar
Heo, Taekyung
Kanter, David
Ladd, Josh
Li, Ziwei
Liu, Winston
Man, Changhai
Mihailescu, Dan
More, Spandan
Park, Joongun
Ramachandran, Ashwin
Ramakrishnaiah, Vinay
Rashidi, Saeed
Reddi, Vijay Janapa
Sharma, Puneet
Tian, Phio
Won, William
Wu, Hanjiang
Xu, Huan
Yoo, Jinsun
Krishna, Tushar
author_facet Sridharan, Srinivas
Badea, Theodor-Adrian
Balogh, Andy
Beckmann, Bradford M.
Coutinho, Brian
Feng, Louis
Fu, Sheng
Gao, Sanshan
Garakani, Mehryar
Heo, Taekyung
Kanter, David
Ladd, Josh
Li, Ziwei
Liu, Winston
Man, Changhai
Mihailescu, Dan
More, Spandan
Park, Joongun
Ramachandran, Ashwin
Ramakrishnaiah, Vinay
Rashidi, Saeed
Reddi, Vijay Janapa
Sharma, Puneet
Tian, Phio
Won, William
Wu, Hanjiang
Xu, Huan
Yoo, Jinsun
Krishna, Tushar
contents The fast pace of artificial intelligence~(AI) innovation demands an agile methodology for observation, reproduction and optimization of distributed machine learning~(ML) workload behavior in production AI systems and enables efficient software-hardware~(SW-HW) co-design for future systems. We present Chakra, an open and portable ecosystem for performance benchmarking and co-design. The core component of Chakra is an open and interoperable graph-based representation of distributed AI/ML workloads, called Chakra execution trace~(ET). These ETs represent key operations, such as compute, memory, and communication, data and control dependencies, timing, and resource constraints. Additionally, Chakra includes a complementary set of tools and capabilities to enable the collection, analysis, generation, and adoption of Chakra ETs by a broad range of simulators, emulators, and replay tools. We present analysis of Chakra ETs collected on production AI clusters and demonstrate value via real-world case studies. Chakra has been adopted by MLCommons and has active contributions and engagement across the industry, including but not limited to NVIDIA, AMD, Meta, Keysight, HPE, and Scala, to name a few.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11333
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces
Sridharan, Srinivas
Badea, Theodor-Adrian
Balogh, Andy
Beckmann, Bradford M.
Coutinho, Brian
Feng, Louis
Fu, Sheng
Gao, Sanshan
Garakani, Mehryar
Heo, Taekyung
Kanter, David
Ladd, Josh
Li, Ziwei
Liu, Winston
Man, Changhai
Mihailescu, Dan
More, Spandan
Park, Joongun
Ramachandran, Ashwin
Ramakrishnaiah, Vinay
Rashidi, Saeed
Reddi, Vijay Janapa
Sharma, Puneet
Tian, Phio
Won, William
Wu, Hanjiang
Xu, Huan
Yoo, Jinsun
Krishna, Tushar
Distributed, Parallel, and Cluster Computing
Machine Learning
Performance
The fast pace of artificial intelligence~(AI) innovation demands an agile methodology for observation, reproduction and optimization of distributed machine learning~(ML) workload behavior in production AI systems and enables efficient software-hardware~(SW-HW) co-design for future systems. We present Chakra, an open and portable ecosystem for performance benchmarking and co-design. The core component of Chakra is an open and interoperable graph-based representation of distributed AI/ML workloads, called Chakra execution trace~(ET). These ETs represent key operations, such as compute, memory, and communication, data and control dependencies, timing, and resource constraints. Additionally, Chakra includes a complementary set of tools and capabilities to enable the collection, analysis, generation, and adoption of Chakra ETs by a broad range of simulators, emulators, and replay tools. We present analysis of Chakra ETs collected on production AI clusters and demonstrate value via real-world case studies. Chakra has been adopted by MLCommons and has active contributions and engagement across the industry, including but not limited to NVIDIA, AMD, Meta, Keysight, HPE, and Scala, to name a few.
title MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces
topic Distributed, Parallel, and Cluster Computing
Machine Learning
Performance
url https://arxiv.org/abs/2605.11333