MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913143251271680 |
|---|---|
| author | Sridharan, Srinivas Badea, Theodor-Adrian Balogh, Andy Beckmann, Bradford M. Coutinho, Brian Feng, Louis Fu, Sheng Gao, Sanshan Garakani, Mehryar Heo, Taekyung Kanter, David Ladd, Josh Li, Ziwei Liu, Winston Man, Changhai Mihailescu, Dan More, Spandan Park, Joongun Ramachandran, Ashwin Ramakrishnaiah, Vinay Rashidi, Saeed Reddi, Vijay Janapa Sharma, Puneet Tian, Phio Won, William Wu, Hanjiang Xu, Huan Yoo, Jinsun Krishna, Tushar |
| author_facet | Sridharan, Srinivas Badea, Theodor-Adrian Balogh, Andy Beckmann, Bradford M. Coutinho, Brian Feng, Louis Fu, Sheng Gao, Sanshan Garakani, Mehryar Heo, Taekyung Kanter, David Ladd, Josh Li, Ziwei Liu, Winston Man, Changhai Mihailescu, Dan More, Spandan Park, Joongun Ramachandran, Ashwin Ramakrishnaiah, Vinay Rashidi, Saeed Reddi, Vijay Janapa Sharma, Puneet Tian, Phio Won, William Wu, Hanjiang Xu, Huan Yoo, Jinsun Krishna, Tushar |
| contents | The fast pace of artificial intelligence~(AI) innovation demands an agile methodology for observation, reproduction and optimization of distributed machine learning~(ML) workload behavior in production AI systems and enables efficient software-hardware~(SW-HW) co-design for future systems. We present Chakra, an open and portable ecosystem for performance benchmarking and co-design. The core component of Chakra is an open and interoperable graph-based representation of distributed AI/ML workloads, called Chakra execution trace~(ET). These ETs represent key operations, such as compute, memory, and communication, data and control dependencies, timing, and resource constraints. Additionally, Chakra includes a complementary set of tools and capabilities to enable the collection, analysis, generation, and adoption of Chakra ETs by a broad range of simulators, emulators, and replay tools. We present analysis of Chakra ETs collected on production AI clusters and demonstrate value via real-world case studies. Chakra has been adopted by MLCommons and has active contributions and engagement across the industry, including but not limited to NVIDIA, AMD, Meta, Keysight, HPE, and Scala, to name a few. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_11333 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces Sridharan, Srinivas Badea, Theodor-Adrian Balogh, Andy Beckmann, Bradford M. Coutinho, Brian Feng, Louis Fu, Sheng Gao, Sanshan Garakani, Mehryar Heo, Taekyung Kanter, David Ladd, Josh Li, Ziwei Liu, Winston Man, Changhai Mihailescu, Dan More, Spandan Park, Joongun Ramachandran, Ashwin Ramakrishnaiah, Vinay Rashidi, Saeed Reddi, Vijay Janapa Sharma, Puneet Tian, Phio Won, William Wu, Hanjiang Xu, Huan Yoo, Jinsun Krishna, Tushar Distributed, Parallel, and Cluster Computing Machine Learning Performance The fast pace of artificial intelligence~(AI) innovation demands an agile methodology for observation, reproduction and optimization of distributed machine learning~(ML) workload behavior in production AI systems and enables efficient software-hardware~(SW-HW) co-design for future systems. We present Chakra, an open and portable ecosystem for performance benchmarking and co-design. The core component of Chakra is an open and interoperable graph-based representation of distributed AI/ML workloads, called Chakra execution trace~(ET). These ETs represent key operations, such as compute, memory, and communication, data and control dependencies, timing, and resource constraints. Additionally, Chakra includes a complementary set of tools and capabilities to enable the collection, analysis, generation, and adoption of Chakra ETs by a broad range of simulators, emulators, and replay tools. We present analysis of Chakra ETs collected on production AI clusters and demonstrate value via real-world case studies. Chakra has been adopted by MLCommons and has active contributions and engagement across the industry, including but not limited to NVIDIA, AMD, Meta, Keysight, HPE, and Scala, to name a few. |
| title | MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces |
| topic | Distributed, Parallel, and Cluster Computing Machine Learning Performance |
| url | https://arxiv.org/abs/2605.11333 |