SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Kuan, Zhang, Shuo, Wang, Huacan, Yu, Fangzhou, Sheng, Zecheng, Gu, Yi, Ming, Weipeng, Xue, Lei, Liu, Chen, Hu, Sen, Chen, Ronghao, Lin, Siyue, Hou, Yuqing, Mou, Xiaofeng, Xu, Yi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911739650506752
author Li, Kuan
Zhang, Shuo
Wang, Huacan
Yu, Fangzhou
Sheng, Zecheng
Gu, Yi
Ming, Weipeng
Xue, Lei
Liu, Chen
Hu, Sen
Chen, Ronghao
Lin, Siyue
Hou, Yuqing
Mou, Xiaofeng
Xu, Yi
author_facet Li, Kuan
Zhang, Shuo
Wang, Huacan
Yu, Fangzhou
Sheng, Zecheng
Gu, Yi
Ming, Weipeng
Xue, Lei
Liu, Chen
Hu, Sen
Chen, Ronghao
Lin, Siyue
Hou, Yuqing
Mou, Xiaofeng
Xu, Yi
contents Smart homes are evolving toward complex state-dependent living environments, requiring Large Language Models (LLMs) to reason over user intent, preferences, and multi-device interactions. However, existing smart-home benchmarks often focus on static instruction-to-API mapping or limited simulations, failing to evaluate whether LLMs can reason, interact, and act reliably in realistic household scenarios. To address these limitations, we introduce SMH-Bench, a comprehensive benchmark for evaluating LLMs in smart-home environments. Built upon HomeEnv, an executable and verifiable smart-home simulator, SMH-Bench contains 1,100 high-quality tasks spanning 7 categories and 22 fine-grained subcategories. It further stratifies tasks across simple, medium and complex homes, ranging from small apartments to dense multi-room environments with 135 devices. Experiments show that although frontier LLMs achieve strong performance on explicit control and query tasks, they still exhibit significant weaknesses in automation task scheduling, ambiguity handling and personalized reasoning, especially as home complexity increases. We hope SMH-Bench will facilitate the development of more reliable, context-aware, and practically deployable smart-home agents.
format Preprint
id arxiv_https___arxiv_org_abs_2606_01912
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes
Li, Kuan
Zhang, Shuo
Wang, Huacan
Yu, Fangzhou
Sheng, Zecheng
Gu, Yi
Ming, Weipeng
Xue, Lei
Liu, Chen
Hu, Sen
Chen, Ronghao
Lin, Siyue
Hou, Yuqing
Mou, Xiaofeng
Xu, Yi
Artificial Intelligence
Smart homes are evolving toward complex state-dependent living environments, requiring Large Language Models (LLMs) to reason over user intent, preferences, and multi-device interactions. However, existing smart-home benchmarks often focus on static instruction-to-API mapping or limited simulations, failing to evaluate whether LLMs can reason, interact, and act reliably in realistic household scenarios. To address these limitations, we introduce SMH-Bench, a comprehensive benchmark for evaluating LLMs in smart-home environments. Built upon HomeEnv, an executable and verifiable smart-home simulator, SMH-Bench contains 1,100 high-quality tasks spanning 7 categories and 22 fine-grained subcategories. It further stratifies tasks across simple, medium and complex homes, ranging from small apartments to dense multi-room environments with 135 devices. Experiments show that although frontier LLMs achieve strong performance on explicit control and query tasks, they still exhibit significant weaknesses in automation task scheduling, ambiguity handling and personalized reasoning, especially as home complexity increases. We hope SMH-Bench will facilitate the development of more reliable, context-aware, and practically deployable smart-home agents.
title SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes
topic Artificial Intelligence
url https://arxiv.org/abs/2606.01912