QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Prakash, Shvetank, Cheng, Andrew, Tschand, Arya, Mazumder, Mark, Gohil, Varun, Ma, Jeffrey, Yik, Jason, Wan, Zishen, Quaye, Jessica, Alvanaki, Elisavet Lydia, Kumar, Avinash, Mazumdar, Chandrashis, Khare, Tuhin, Ingare, Alexander, Uchendu, Ikechukwu, Ghosal, Radhika, Tyagi, Abhishek, Wang, Chenyu, Garavagno, Andrea Mattia, Gu, Sarah, Guo, Alice, Hur, Grace, Carloni, Luca, Krishna, Tushar, Nayak, Ankita, Yazdanbakhsh, Amir, Reddi, Vijay Janapa
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917042105352192
author Prakash, Shvetank
Cheng, Andrew
Tschand, Arya
Mazumder, Mark
Gohil, Varun
Ma, Jeffrey
Yik, Jason
Wan, Zishen
Quaye, Jessica
Alvanaki, Elisavet Lydia
Kumar, Avinash
Mazumdar, Chandrashis
Khare, Tuhin
Ingare, Alexander
Uchendu, Ikechukwu
Ghosal, Radhika
Tyagi, Abhishek
Wang, Chenyu
Garavagno, Andrea Mattia
Gu, Sarah
Guo, Alice
Hur, Grace
Carloni, Luca
Krishna, Tushar
Nayak, Ankita
Yazdanbakhsh, Amir
Reddi, Vijay Janapa
author_facet Prakash, Shvetank
Cheng, Andrew
Tschand, Arya
Mazumder, Mark
Gohil, Varun
Ma, Jeffrey
Yik, Jason
Wan, Zishen
Quaye, Jessica
Alvanaki, Elisavet Lydia
Kumar, Avinash
Mazumdar, Chandrashis
Khare, Tuhin
Ingare, Alexander
Uchendu, Ikechukwu
Ghosal, Radhika
Tyagi, Abhishek
Wang, Chenyu
Garavagno, Andrea Mattia
Gu, Sarah
Guo, Alice
Hur, Grace
Carloni, Luca
Krishna, Tushar
Nayak, Ankita
Yazdanbakhsh, Amir
Reddi, Vijay Janapa
contents The field of computer architecture, which bridges high-level software abstractions and low-level hardware implementations, remains absent from current large language model (LLM) evaluations. To this end, we present QuArch (pronounced 'quark'), the first benchmark designed to facilitate the development and evaluation of LLM knowledge and reasoning capabilities specifically in computer architecture. QuArch provides a comprehensive collection of 2,671 expert-validated question-answer (QA) pairs covering various aspects of computer architecture, including processor design, memory systems, and interconnection networks. Our evaluation reveals that while frontier models possess domain-specific knowledge, they struggle with skills that require higher-order thinking in computer architecture. Frontier model accuracies vary widely (from 34% to 72%) on these advanced questions, highlighting persistent gaps in architectural reasoning across analysis, design, and implementation QAs. By holistically assessing fundamental skills, QuArch provides a foundation for building and measuring LLM capabilities that can accelerate innovation in computing systems. With over 140 contributors from 40 institutions, this benchmark represents a community effort to set the standard for architectural reasoning in LLM evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22087
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture
Prakash, Shvetank
Cheng, Andrew
Tschand, Arya
Mazumder, Mark
Gohil, Varun
Ma, Jeffrey
Yik, Jason
Wan, Zishen
Quaye, Jessica
Alvanaki, Elisavet Lydia
Kumar, Avinash
Mazumdar, Chandrashis
Khare, Tuhin
Ingare, Alexander
Uchendu, Ikechukwu
Ghosal, Radhika
Tyagi, Abhishek
Wang, Chenyu
Garavagno, Andrea Mattia
Gu, Sarah
Guo, Alice
Hur, Grace
Carloni, Luca
Krishna, Tushar
Nayak, Ankita
Yazdanbakhsh, Amir
Reddi, Vijay Janapa
Hardware Architecture
Artificial Intelligence
Machine Learning
Software Engineering
The field of computer architecture, which bridges high-level software abstractions and low-level hardware implementations, remains absent from current large language model (LLM) evaluations. To this end, we present QuArch (pronounced 'quark'), the first benchmark designed to facilitate the development and evaluation of LLM knowledge and reasoning capabilities specifically in computer architecture. QuArch provides a comprehensive collection of 2,671 expert-validated question-answer (QA) pairs covering various aspects of computer architecture, including processor design, memory systems, and interconnection networks. Our evaluation reveals that while frontier models possess domain-specific knowledge, they struggle with skills that require higher-order thinking in computer architecture. Frontier model accuracies vary widely (from 34% to 72%) on these advanced questions, highlighting persistent gaps in architectural reasoning across analysis, design, and implementation QAs. By holistically assessing fundamental skills, QuArch provides a foundation for building and measuring LLM capabilities that can accelerate innovation in computing systems. With over 140 contributors from 40 institutions, this benchmark represents a community effort to set the standard for architectural reasoning in LLM evaluation.
title QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture
topic Hardware Architecture
Artificial Intelligence
Machine Learning
Software Engineering
url https://arxiv.org/abs/2510.22087