FactoryBench: Evaluating Industrial Machine Understanding

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Merzouki, Yanis, Izquierdo, Coral, Ignuta-Ciuncanu, Matei, Gomez-Bracamonte, Marcos, Maggioni, Riccardo, Lombardi, Alessandro, Mazzoleni, Camilla, Martelli, Federico, Gunther, Balazs, Petersen, Jonas, Petersen, Philipp
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915992592973824
author Merzouki, Yanis
Izquierdo, Coral
Ignuta-Ciuncanu, Matei
Gomez-Bracamonte, Marcos
Maggioni, Riccardo
Lombardi, Alessandro
Mazzoleni, Camilla
Martelli, Federico
Gunther, Balazs
Petersen, Jonas
Petersen, Philipp
author_facet Merzouki, Yanis
Izquierdo, Coral
Ignuta-Ciuncanu, Matei
Gomez-Bracamonte, Marcos
Maggioni, Riccardo
Lombardi, Alessandro
Mazzoleni, Camilla
Martelli, Federico
Gunther, Balazs
Petersen, Jonas
Petersen, Philipp
contents We introduce FactoryBench, a benchmark for evaluating time-series models and LLMs on machine understanding over industrial robotic telemetry. Q&A pairs are organized along four causal levels (state, intervention, counterfactual, decision) instantiating Pearl's ladder of causation, and span five answer formats: four structured formats are scored deterministically and free-form answers are scored by an LLM-as-judge voting protocol. We propose a scalable Q&A generation framework built around structured question templates, present FactoryWave (a dense, multitask, multivariate sensor dataset collected from a UR3 cobot and a KUKA KR10 industrial arm), and construct FactoryBench as a large-scale benchmark of over 70k Q&A items grounded in roughly 15k normalized episodes from FactoryWave, AURSAD, and voraus-AD. Zero-shot evaluation of six frontier LLMs shows that no model exceeds 50% on structured levels or 18% on decision-making, revealing a wide gap between current models and operational machine understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2605_07675
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FactoryBench: Evaluating Industrial Machine Understanding
Merzouki, Yanis
Izquierdo, Coral
Ignuta-Ciuncanu, Matei
Gomez-Bracamonte, Marcos
Maggioni, Riccardo
Lombardi, Alessandro
Mazzoleni, Camilla
Martelli, Federico
Gunther, Balazs
Petersen, Jonas
Petersen, Philipp
Artificial Intelligence
Machine Learning
cs.AI (primary), cs.RO, cs.LG, cs.CL
I.2.6; I.2.9
We introduce FactoryBench, a benchmark for evaluating time-series models and LLMs on machine understanding over industrial robotic telemetry. Q&A pairs are organized along four causal levels (state, intervention, counterfactual, decision) instantiating Pearl's ladder of causation, and span five answer formats: four structured formats are scored deterministically and free-form answers are scored by an LLM-as-judge voting protocol. We propose a scalable Q&A generation framework built around structured question templates, present FactoryWave (a dense, multitask, multivariate sensor dataset collected from a UR3 cobot and a KUKA KR10 industrial arm), and construct FactoryBench as a large-scale benchmark of over 70k Q&A items grounded in roughly 15k normalized episodes from FactoryWave, AURSAD, and voraus-AD. Zero-shot evaluation of six frontier LLMs shows that no model exceeds 50% on structured levels or 18% on decision-making, revealing a wide gap between current models and operational machine understanding.
title FactoryBench: Evaluating Industrial Machine Understanding
topic Artificial Intelligence
Machine Learning
cs.AI (primary), cs.RO, cs.LG, cs.CL
I.2.6; I.2.9
url https://arxiv.org/abs/2605.07675