SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Seo, Gyuhyeon, Yang, Jungwoo, Pyo, Junseong, Kim, Nalim, Lee, Jonggeun, Jo, Yohan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918365157654528
author Seo, Gyuhyeon
Yang, Jungwoo
Pyo, Junseong
Kim, Nalim
Lee, Jonggeun
Jo, Yohan
author_facet Seo, Gyuhyeon
Yang, Jungwoo
Pyo, Junseong
Kim, Nalim
Lee, Jonggeun
Jo, Yohan
contents We introduce $\textbf{SimuHome}$, a high-fidelity smart home simulator and a benchmark of 600 episodes for LLM-based smart home agents. Existing smart home benchmarks treat the home as a static system, neither simulating how device operations affect environmental variables over time nor supporting workflow scheduling of device commands. SimuHome is grounded in the Matter protocol, the industry standard that defines how real smart home devices communicate and operate. Agents interact with devices through SimuHome's APIs and observe how their actions continuously affect environmental variables such as temperature and humidity. Our benchmark covers state inquiry, implicit user intent inference, explicit device control, and workflow scheduling, each with both feasible and infeasible requests. For workflow scheduling, the simulator accelerates time so that scheduled workflows can be evaluated immediately. An evaluation of 18 agents reveals that workflow scheduling is the hardest category, with failures persisting across alternative agent frameworks and fine-tuning. These findings suggest that SimuHome's time-accelerated simulation could serve as an environment for agents to pre-validate their actions before committing them to the real world.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24282
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
Seo, Gyuhyeon
Yang, Jungwoo
Pyo, Junseong
Kim, Nalim
Lee, Jonggeun
Jo, Yohan
Computation and Language
Artificial Intelligence
We introduce $\textbf{SimuHome}$, a high-fidelity smart home simulator and a benchmark of 600 episodes for LLM-based smart home agents. Existing smart home benchmarks treat the home as a static system, neither simulating how device operations affect environmental variables over time nor supporting workflow scheduling of device commands. SimuHome is grounded in the Matter protocol, the industry standard that defines how real smart home devices communicate and operate. Agents interact with devices through SimuHome's APIs and observe how their actions continuously affect environmental variables such as temperature and humidity. Our benchmark covers state inquiry, implicit user intent inference, explicit device control, and workflow scheduling, each with both feasible and infeasible requests. For workflow scheduling, the simulator accelerates time so that scheduled workflows can be evaluated immediately. An evaluation of 18 agents reveals that workflow scheduling is the hardest category, with failures persisting across alternative agent frameworks and fine-tuning. These findings suggest that SimuHome's time-accelerated simulation could serve as an environment for agents to pre-validate their actions before committing them to the real world.
title SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.24282