HomeGuard: VLM-based Embodied Safeguard for Identifying Contextual Risk in Household Task

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lu, Xiaoya, Zhou, Yijin, Chen, Zeren, Wang, Ruocheng, Sima, Bingrui, Zhou, Enshen, Sheng, Lu, Liu, Dongrui, Shao, Jing
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917344685588480
author Lu, Xiaoya
Zhou, Yijin
Chen, Zeren
Wang, Ruocheng
Sima, Bingrui
Zhou, Enshen
Sheng, Lu
Liu, Dongrui
Shao, Jing
author_facet Lu, Xiaoya
Zhou, Yijin
Chen, Zeren
Wang, Ruocheng
Sima, Bingrui
Zhou, Enshen
Sheng, Lu
Liu, Dongrui
Shao, Jing
contents Vision-Language Models (VLMs) empower embodied agents to execute complex instructions, yet they remain vulnerable to contextual safety risks where benign commands become hazardous due to subtle environmental states. Existing safeguards often prove inadequate. Rule-based methods lack scalability in object-dense scenes, whereas model-based approaches relying on prompt engineering suffer from unfocused perception, resulting in missed risks or hallucinations. To address this, we propose an architecture-agnostic safeguard featuring Context-Guided Chain-of-Thought (CG-CoT). This mechanism decomposes risk assessment into active perception that sequentially anchors attention to interaction targets and relevant spatial neighborhoods, followed by semantic judgment based on this visual evidence. We support this approach with a curated grounding dataset and a two-stage training strategy utilizing Reinforcement Fine-Tuning (RFT) with process rewards to enforce precise intermediate grounding. Experiments demonstrate that our model HomeGuard significantly enhances safety, improving risk match rates by over 30% compared to base models while reducing oversafety. Beyond hazard detection, the generated visual anchors serve as actionable spatial constraints for downstream planners, facilitating explicit collision avoidance and safety trajectory generation. Code and data are released under https://github.com/AI45Lab/HomeGuard
format Preprint
id arxiv_https___arxiv_org_abs_2603_14367
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HomeGuard: VLM-based Embodied Safeguard for Identifying Contextual Risk in Household Task
Lu, Xiaoya
Zhou, Yijin
Chen, Zeren
Wang, Ruocheng
Sima, Bingrui
Zhou, Enshen
Sheng, Lu
Liu, Dongrui
Shao, Jing
Computer Vision and Pattern Recognition
Vision-Language Models (VLMs) empower embodied agents to execute complex instructions, yet they remain vulnerable to contextual safety risks where benign commands become hazardous due to subtle environmental states. Existing safeguards often prove inadequate. Rule-based methods lack scalability in object-dense scenes, whereas model-based approaches relying on prompt engineering suffer from unfocused perception, resulting in missed risks or hallucinations. To address this, we propose an architecture-agnostic safeguard featuring Context-Guided Chain-of-Thought (CG-CoT). This mechanism decomposes risk assessment into active perception that sequentially anchors attention to interaction targets and relevant spatial neighborhoods, followed by semantic judgment based on this visual evidence. We support this approach with a curated grounding dataset and a two-stage training strategy utilizing Reinforcement Fine-Tuning (RFT) with process rewards to enforce precise intermediate grounding. Experiments demonstrate that our model HomeGuard significantly enhances safety, improving risk match rates by over 30% compared to base models while reducing oversafety. Beyond hazard detection, the generated visual anchors serve as actionable spatial constraints for downstream planners, facilitating explicit collision avoidance and safety trajectory generation. Code and data are released under https://github.com/AI45Lab/HomeGuard
title HomeGuard: VLM-based Embodied Safeguard for Identifying Contextual Risk in Household Task
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.14367