Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866911252844904448 |
|---|---|
| author | Guo, Xingang Li, Yaxin Kong, Xiangyi Jiang, Yilan Zhao, Xiayu Gong, Zhihua Zhang, Yufan Li, Daixuan Sang, Tianle Zhu, Beixiao Jun, Gregory Huang, Yingbing Liu, Yiqi Xue, Yuqi Kundu, Rahul Dev Lim, Qi Jian Zhao, Yizhou Granger, Luke Alexander Younis, Mohamed Badr Keivan, Darioush Sabharwal, Nippun Sinha, Shreyanka Agarwal, Prakhar Vandyck, Kojo Mai, Hanlin Wang, Zichen Venkatesh, Aditya Barik, Ayush Yang, Jiankun Yue, Chongying He, Jingjie Wang, Libin Xu, Licheng Chen, Hao Wang, Jinwen Xu, Liujun Shetty, Rushabh Guo, Ziheng Song, Dahui Jha, Manvi Liang, Weijie Yan, Weiman Zhang, Bryan Karnoor, Sahil Bhandary Zhang, Jialiang Pandya, Rutva Gong, Xinyi Ganesh, Mithesh Ballae Shi, Feize Xu, Ruiling Zhang, Yifan Ouyang, Yanfeng Qin, Lianhui Rosenbaum, Elyse Snyder, Corey Seiler, Peter Dullerud, Geir Zhang, Xiaojia Shelly Cheng, Zuofu Hanumolu, Pavan Kumar Huang, Jian Kulkarni, Mayank Namazifar, Mahdi Zhang, Huan Hu, Bin |
| author_facet | Guo, Xingang Li, Yaxin Kong, Xiangyi Jiang, Yilan Zhao, Xiayu Gong, Zhihua Zhang, Yufan Li, Daixuan Sang, Tianle Zhu, Beixiao Jun, Gregory Huang, Yingbing Liu, Yiqi Xue, Yuqi Kundu, Rahul Dev Lim, Qi Jian Zhao, Yizhou Granger, Luke Alexander Younis, Mohamed Badr Keivan, Darioush Sabharwal, Nippun Sinha, Shreyanka Agarwal, Prakhar Vandyck, Kojo Mai, Hanlin Wang, Zichen Venkatesh, Aditya Barik, Ayush Yang, Jiankun Yue, Chongying He, Jingjie Wang, Libin Xu, Licheng Chen, Hao Wang, Jinwen Xu, Liujun Shetty, Rushabh Guo, Ziheng Song, Dahui Jha, Manvi Liang, Weijie Yan, Weiman Zhang, Bryan Karnoor, Sahil Bhandary Zhang, Jialiang Pandya, Rutva Gong, Xinyi Ganesh, Mithesh Ballae Shi, Feize Xu, Ruiling Zhang, Yifan Ouyang, Yanfeng Qin, Lianhui Rosenbaum, Elyse Snyder, Corey Seiler, Peter Dullerud, Geir Zhang, Xiaojia Shelly Cheng, Zuofu Hanumolu, Pavan Kumar Huang, Jian Kulkarni, Mayank Namazifar, Mahdi Zhang, Huan Hu, Bin |
| contents | Modern engineering, spanning electrical, mechanical, aerospace, civil, and computer disciplines, stands as a cornerstone of human civilization and the foundation of our society. However, engineering design poses a fundamentally different challenge for large language models (LLMs) compared with traditional textbook-style problem solving or factual question answering. Although existing benchmarks have driven progress in areas such as language understanding, code synthesis, and scientific problem solving, real-world engineering design demands the synthesis of domain knowledge, navigation of complex trade-offs, and management of the tedious processes that consume much of practicing engineers' time. Despite these shared challenges across engineering disciplines, no benchmark currently captures the unique demands of engineering design work. In this work, we introduce EngDesign, an Engineering Design benchmark that evaluates LLMs' abilities to perform practical design tasks across nine engineering domains. Unlike existing benchmarks that focus on factual recall or question answering, EngDesign uniquely emphasizes LLMs' ability to synthesize domain knowledge, reason under constraints, and generate functional, objective-oriented engineering designs. Each task in EngDesign represents a real-world engineering design problem, accompanied by a detailed task description specifying design goals, constraints, and performance requirements. EngDesign pioneers a simulation-based evaluation paradigm that moves beyond textbook knowledge to assess genuine engineering design capabilities and shifts evaluation from static answer checking to dynamic, simulation-driven functional verification, marking a crucial step toward realizing the vision of engineering Artificial General Intelligence (AGI). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_16204 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs Guo, Xingang Li, Yaxin Kong, Xiangyi Jiang, Yilan Zhao, Xiayu Gong, Zhihua Zhang, Yufan Li, Daixuan Sang, Tianle Zhu, Beixiao Jun, Gregory Huang, Yingbing Liu, Yiqi Xue, Yuqi Kundu, Rahul Dev Lim, Qi Jian Zhao, Yizhou Granger, Luke Alexander Younis, Mohamed Badr Keivan, Darioush Sabharwal, Nippun Sinha, Shreyanka Agarwal, Prakhar Vandyck, Kojo Mai, Hanlin Wang, Zichen Venkatesh, Aditya Barik, Ayush Yang, Jiankun Yue, Chongying He, Jingjie Wang, Libin Xu, Licheng Chen, Hao Wang, Jinwen Xu, Liujun Shetty, Rushabh Guo, Ziheng Song, Dahui Jha, Manvi Liang, Weijie Yan, Weiman Zhang, Bryan Karnoor, Sahil Bhandary Zhang, Jialiang Pandya, Rutva Gong, Xinyi Ganesh, Mithesh Ballae Shi, Feize Xu, Ruiling Zhang, Yifan Ouyang, Yanfeng Qin, Lianhui Rosenbaum, Elyse Snyder, Corey Seiler, Peter Dullerud, Geir Zhang, Xiaojia Shelly Cheng, Zuofu Hanumolu, Pavan Kumar Huang, Jian Kulkarni, Mayank Namazifar, Mahdi Zhang, Huan Hu, Bin Computational Engineering, Finance, and Science Human-Computer Interaction Robotics Modern engineering, spanning electrical, mechanical, aerospace, civil, and computer disciplines, stands as a cornerstone of human civilization and the foundation of our society. However, engineering design poses a fundamentally different challenge for large language models (LLMs) compared with traditional textbook-style problem solving or factual question answering. Although existing benchmarks have driven progress in areas such as language understanding, code synthesis, and scientific problem solving, real-world engineering design demands the synthesis of domain knowledge, navigation of complex trade-offs, and management of the tedious processes that consume much of practicing engineers' time. Despite these shared challenges across engineering disciplines, no benchmark currently captures the unique demands of engineering design work. In this work, we introduce EngDesign, an Engineering Design benchmark that evaluates LLMs' abilities to perform practical design tasks across nine engineering domains. Unlike existing benchmarks that focus on factual recall or question answering, EngDesign uniquely emphasizes LLMs' ability to synthesize domain knowledge, reason under constraints, and generate functional, objective-oriented engineering designs. Each task in EngDesign represents a real-world engineering design problem, accompanied by a detailed task description specifying design goals, constraints, and performance requirements. EngDesign pioneers a simulation-based evaluation paradigm that moves beyond textbook knowledge to assess genuine engineering design capabilities and shifts evaluation from static answer checking to dynamic, simulation-driven functional verification, marking a crucial step toward realizing the vision of engineering Artificial General Intelligence (AGI). |
| title | Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs |
| topic | Computational Engineering, Finance, and Science Human-Computer Interaction Robotics |
| url | https://arxiv.org/abs/2509.16204 |