SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Kuan, Zhang, Shuo, Wang, Huacan, Yu, Fangzhou, Sheng, Zecheng, Gu, Yi, Ming, Weipeng, Xue, Lei, Liu, Chen, Hu, Sen, Chen, Ronghao, Lin, Siyue, Hou, Yuqing, Mou, Xiaofeng, Xu, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HomeFlow: A Data Flywheel for Smart Home Agent Training with Verifiable Simulation
by: Gu, Yi, et al.
Published: (2026)
by: Gu, Yi, et al.
Published: (2026)
SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agents
by: Zhou, Yifan, et al.
Published: (2026)
by: Zhou, Yifan, et al.
Published: (2026)
KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions
by: Wu, Tingyu, et al.
Published: (2026)
by: Wu, Tingyu, et al.
Published: (2026)
Reject or Not?: A Benchmark for Voice Assistant Query Rejection in Smart Home Scenario and an Improved Method Based on LLMs
by: Men, Huichao, et al.
Published: (2025)
by: Men, Huichao, et al.
Published: (2025)
CloneMem: Benchmarking Long-Term Memory for AI Clones
by: Hu, Sen, et al.
Published: (2026)
by: Hu, Sen, et al.
Published: (2026)
FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments
by: Yang, Zhi, et al.
Published: (2026)
by: Yang, Zhi, et al.
Published: (2026)
Controlled Self-Evolution for Algorithmic Code Optimization
by: Hu, Tu, et al.
Published: (2026)
by: Hu, Tu, et al.
Published: (2026)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
by: Ni, Ziyi, et al.
Published: (2025)
by: Ni, Ziyi, et al.
Published: (2025)
SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Large Language Models
by: Zhao, Xinyi, et al.
Published: (2025)
by: Zhao, Xinyi, et al.
Published: (2025)
SemaClaw: A Step Towards General-Purpose Personal AI Agents through Harness Engineering
by: Zhu, Ningyan, et al.
Published: (2026)
by: Zhu, Ningyan, et al.
Published: (2026)
Sema Code: Decoupling AI Coding Agents into Programmable, Embeddable Infrastructure
by: Wang, Huacan, et al.
Published: (2026)
by: Wang, Huacan, et al.
Published: (2026)
RealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction
by: Bian, Haonan, et al.
Published: (2026)
by: Bian, Haonan, et al.
Published: (2026)
On the pro-étale cohomology of quotient stacks of Drinfeld spaces
by: Yi, Zecheng
Published: (2025)
by: Yi, Zecheng
Published: (2025)
SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
by: Seo, Gyuhyeon, et al.
Published: (2025)
by: Seo, Gyuhyeon, et al.
Published: (2025)
PersonalHomeBench: Evaluating Agents in Personalized Smart Homes
by: Bharadwaj, Manasa, et al.
Published: (2026)
by: Bharadwaj, Manasa, et al.
Published: (2026)
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation
by: Liu, Yi, et al.
Published: (2026)
by: Liu, Yi, et al.
Published: (2026)
Does Memory Need Graphs? A Unified Framework and Empirical Analysis for Long-Term Dialog Memory
by: Hu, Sen, et al.
Published: (2026)
by: Hu, Sen, et al.
Published: (2026)
MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs
by: Xu, Yunqiu, et al.
Published: (2024)
by: Xu, Yunqiu, et al.
Published: (2024)
End-to-End Direction-Aware Keyword Spotting with Spatial Priors in Noisy Environments
by: Wang, Rui, et al.
Published: (2026)
by: Wang, Rui, et al.
Published: (2026)
UIPress: Bringing Optical Token Compression to UI-to-Code Generation
by: Dai, Dasen, et al.
Published: (2026)
by: Dai, Dasen, et al.
Published: (2026)
VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
Prototype Smart Home Environment With Biofeedback
by: Kamal, Azmyin Md., et al.
Published: (2024)
by: Kamal, Azmyin Md., et al.
Published: (2024)
Rethinking Token Pruning for Historical Screenshots in GUI Visual Agents: Semantic, Spatial, and Temporal Perspectives
by: Li, Daiqiang, et al.
Published: (2026)
by: Li, Daiqiang, et al.
Published: (2026)
SVAG-Bench: A Large-Scale Benchmark for Multi-Instance Spatio-temporal Video Action Grounding
by: Hannan, Tanveer, et al.
Published: (2025)
by: Hannan, Tanveer, et al.
Published: (2025)
Finslerian Succession Model of the Hyperverse (F-SMH): Detailed Revised Edition
by: Ploumis, Ioannis, et al.
Published: (2026)
by: Ploumis, Ioannis, et al.
Published: (2026)
Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
by: Guo, Yifu, et al.
Published: (2025)
by: Guo, Yifu, et al.
Published: (2025)
Uni-NTFM: A Unified Foundation Model for EEG Signal Representation Learning
by: Chen, Zhisheng, et al.
Published: (2025)
by: Chen, Zhisheng, et al.
Published: (2025)
SmartBench: Evaluating LLMs in Smart Homes with Anomalous Device States and Behavioral Contexts
by: Zou, Qingsong, et al.
Published: (2026)
by: Zou, Qingsong, et al.
Published: (2026)
RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving
by: Wang, Huacan, et al.
Published: (2025)
by: Wang, Huacan, et al.
Published: (2025)
Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning
by: Liu, Chengwen, et al.
Published: (2026)
by: Liu, Chengwen, et al.
Published: (2026)
Deep Analysis of Time Series Data for Smart Grid Startup Strategies: A Transformer-LSTM-PSO Model Approach
by: Zhang, Zecheng
Published: (2024)
by: Zhang, Zecheng
Published: (2024)
Trust Your Memory: Verifiable Control of Smart Homes through Reinforcement Learning with Multi-dimensional Rewards
by: Guo, Kai-Yuan, et al.
Published: (2026)
by: Guo, Kai-Yuan, et al.
Published: (2026)
AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation
by: Shi, Wentao, et al.
Published: (2026)
by: Shi, Wentao, et al.
Published: (2026)
AQUA-Bench: Beyond Finding Answers to Knowing When There Are None in Audio Question Answering
by: Kuan, Chun-Yi, et al.
Published: (2026)
by: Kuan, Chun-Yi, et al.
Published: (2026)
HomeBench: Evaluating LLMs in Smart Homes with Valid and Invalid Instructions Across Single and Multiple Devices
by: Li, Silin, et al.
Published: (2025)
by: Li, Silin, et al.
Published: (2025)
PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
by: Selch, Lukas, et al.
Published: (2025)
by: Selch, Lukas, et al.
Published: (2025)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
by: Fang, I-Sheng, et al.
Published: (2025)
by: Fang, I-Sheng, et al.
Published: (2025)
An O(N) quasi-Ewald splitting method for nanoconfined electrostatics
by: Gan, Zecheng, et al.
Published: (2026)
by: Gan, Zecheng, et al.
Published: (2026)
SPA-Bench: A Comprehensive Benchmark for SmartPhone Agent Evaluation
by: Chen, Jingxuan, et al.
Published: (2024)
by: Chen, Jingxuan, et al.
Published: (2024)
GovRelBench:A Benchmark for Government Domain Relevance
by: Wang, Haiquan, et al.
Published: (2025)
by: Wang, Haiquan, et al.
Published: (2025)
Similar Items
-
HomeFlow: A Data Flywheel for Smart Home Agent Training with Verifiable Simulation
by: Gu, Yi, et al.
Published: (2026) -
SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agents
by: Zhou, Yifan, et al.
Published: (2026) -
KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions
by: Wu, Tingyu, et al.
Published: (2026) -
Reject or Not?: A Benchmark for Voice Assistant Query Rejection in Smart Home Scenario and an Improved Method Based on LLMs
by: Men, Huichao, et al.
Published: (2025) -
CloneMem: Benchmarking Long-Term Memory for AI Clones
by: Hu, Sen, et al.
Published: (2026)