Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Yi, Xue, Zeyu, Liu, Mujie, Zhang, Tongqin, Hu, Yan, Zhao, Zhou, Yang, Chenguang, Lu, Zhenyu
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2509.23107
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908613937725440
author Wang, Yi
Xue, Zeyu
Liu, Mujie
Zhang, Tongqin
Hu, Yan
Zhao, Zhou
Yang, Chenguang
Lu, Zhenyu
author_facet Wang, Yi
Xue, Zeyu
Liu, Mujie
Zhang, Tongqin
Hu, Yan
Zhao, Zhou
Yang, Chenguang
Lu, Zhenyu
contents Teleoperation via natural-language reduces operator workload and enhances safety in high-risk or remote settings. However, in dynamic remote scenes, transmission latency during bidirectional communication creates gaps between remote perceived states and operator intent, leading to command misunderstanding and incorrect execution. To mitigate this, we introduce the Spatio-Temporal Open-Vocabulary Scene Graph (ST-OVSG), a representation that enriches open-vocabulary perception with temporal dynamics and lightweight latency annotations. ST-OVSG leverages LVLMs to construct open-vocabulary 3D object representations, and extends them into the temporal domain via Hungarian assignment with our temporal matching cost, yielding a unified spatio-temporal scene graph. A latency tag is embedded to enable LVLM planners to retrospectively query past scene states, thereby resolving local-remote state mismatches caused by transmission delays. To further reduce redundancy and highlight task-relevant cues, we propose a task-oriented subgraph filtering strategy that produces compact inputs for the planner. ST-OVSG generalizes to novel categories and enhances planning robustness against transmission latency without requiring fine-tuning. Experiments show that our method achieves 74 percent node accuracy on the Replica benchmark, outperforming ConceptGraph. Notably, in the latency-robustness experiment, the LVLM planner assisted by ST-OVSG achieved a planning success rate of 70.5 percent.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23107
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Open-Vocabulary Spatio-Temporal Scene Graph for Robot Perception and Teleoperation Planning
Wang, Yi
Xue, Zeyu
Liu, Mujie
Zhang, Tongqin
Hu, Yan
Zhao, Zhou
Yang, Chenguang
Lu, Zhenyu
Robotics
Artificial Intelligence
Teleoperation via natural-language reduces operator workload and enhances safety in high-risk or remote settings. However, in dynamic remote scenes, transmission latency during bidirectional communication creates gaps between remote perceived states and operator intent, leading to command misunderstanding and incorrect execution. To mitigate this, we introduce the Spatio-Temporal Open-Vocabulary Scene Graph (ST-OVSG), a representation that enriches open-vocabulary perception with temporal dynamics and lightweight latency annotations. ST-OVSG leverages LVLMs to construct open-vocabulary 3D object representations, and extends them into the temporal domain via Hungarian assignment with our temporal matching cost, yielding a unified spatio-temporal scene graph. A latency tag is embedded to enable LVLM planners to retrospectively query past scene states, thereby resolving local-remote state mismatches caused by transmission delays. To further reduce redundancy and highlight task-relevant cues, we propose a task-oriented subgraph filtering strategy that produces compact inputs for the planner. ST-OVSG generalizes to novel categories and enhances planning robustness against transmission latency without requiring fine-tuning. Experiments show that our method achieves 74 percent node accuracy on the Replica benchmark, outperforming ConceptGraph. Notably, in the latency-robustness experiment, the LVLM planner assisted by ST-OVSG achieved a planning success rate of 70.5 percent.
title Open-Vocabulary Spatio-Temporal Scene Graph for Robot Perception and Teleoperation Planning
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2509.23107