FactoryBench: Evaluating Industrial Machine Understanding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866915992592973824 |
|---|---|
| author | Merzouki, Yanis Izquierdo, Coral Ignuta-Ciuncanu, Matei Gomez-Bracamonte, Marcos Maggioni, Riccardo Lombardi, Alessandro Mazzoleni, Camilla Martelli, Federico Gunther, Balazs Petersen, Jonas Petersen, Philipp |
| author_facet | Merzouki, Yanis Izquierdo, Coral Ignuta-Ciuncanu, Matei Gomez-Bracamonte, Marcos Maggioni, Riccardo Lombardi, Alessandro Mazzoleni, Camilla Martelli, Federico Gunther, Balazs Petersen, Jonas Petersen, Philipp |
| contents | We introduce FactoryBench, a benchmark for evaluating time-series models and LLMs on machine understanding over industrial robotic telemetry. Q&A pairs are organized along four causal levels (state, intervention, counterfactual, decision) instantiating Pearl's ladder of causation, and span five answer formats: four structured formats are scored deterministically and free-form answers are scored by an LLM-as-judge voting protocol. We propose a scalable Q&A generation framework built around structured question templates, present FactoryWave (a dense, multitask, multivariate sensor dataset collected from a UR3 cobot and a KUKA KR10 industrial arm), and construct FactoryBench as a large-scale benchmark of over 70k Q&A items grounded in roughly 15k normalized episodes from FactoryWave, AURSAD, and voraus-AD. Zero-shot evaluation of six frontier LLMs shows that no model exceeds 50% on structured levels or 18% on decision-making, revealing a wide gap between current models and operational machine understanding. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_07675 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | FactoryBench: Evaluating Industrial Machine Understanding Merzouki, Yanis Izquierdo, Coral Ignuta-Ciuncanu, Matei Gomez-Bracamonte, Marcos Maggioni, Riccardo Lombardi, Alessandro Mazzoleni, Camilla Martelli, Federico Gunther, Balazs Petersen, Jonas Petersen, Philipp Artificial Intelligence Machine Learning cs.AI (primary), cs.RO, cs.LG, cs.CL I.2.6; I.2.9 We introduce FactoryBench, a benchmark for evaluating time-series models and LLMs on machine understanding over industrial robotic telemetry. Q&A pairs are organized along four causal levels (state, intervention, counterfactual, decision) instantiating Pearl's ladder of causation, and span five answer formats: four structured formats are scored deterministically and free-form answers are scored by an LLM-as-judge voting protocol. We propose a scalable Q&A generation framework built around structured question templates, present FactoryWave (a dense, multitask, multivariate sensor dataset collected from a UR3 cobot and a KUKA KR10 industrial arm), and construct FactoryBench as a large-scale benchmark of over 70k Q&A items grounded in roughly 15k normalized episodes from FactoryWave, AURSAD, and voraus-AD. Zero-shot evaluation of six frontier LLMs shows that no model exceeds 50% on structured levels or 18% on decision-making, revealing a wide gap between current models and operational machine understanding. |
| title | FactoryBench: Evaluating Industrial Machine Understanding |
| topic | Artificial Intelligence Machine Learning cs.AI (primary), cs.RO, cs.LG, cs.CL I.2.6; I.2.9 |
| url | https://arxiv.org/abs/2605.07675 |