MERGE: Guided Vision-Language Models for Multi-Actor Event Reasoning and Grounding in Human-Robot Interaction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deigmoeller, Joerg, Agarwal, Nakul, Hasler, Stephan, Tanneberg, Daniel, Belardinelli, Anna, Ghoddoosian, Reza, Wang, Chao, Ocker, Felix, Zhang, Fan, Dariush, Behzad, Gienger, Michael
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917353353117696
author Deigmoeller, Joerg
Agarwal, Nakul
Hasler, Stephan
Tanneberg, Daniel
Belardinelli, Anna
Ghoddoosian, Reza
Wang, Chao
Ocker, Felix
Zhang, Fan
Dariush, Behzad
Gienger, Michael
author_facet Deigmoeller, Joerg
Agarwal, Nakul
Hasler, Stephan
Tanneberg, Daniel
Belardinelli, Anna
Ghoddoosian, Reza
Wang, Chao
Ocker, Felix
Zhang, Fan
Dariush, Behzad
Gienger, Michael
contents We introduce MERGE, a system for situational grounding of actors, objects, and events in dynamic human-robot group interactions. Effective collaboration in such settings requires consistent situational awareness, built on persistent representations of people and objects and an episodic abstraction of events. MERGE achieves this by uniquely identifying physical instances of actors (humans or robots) and objects and structuring them into actor-action-object relations, ensuring temporal consistency across interactions. Central to MERGE is the integration of Vision-Language Models (VLMs) guided with a perception pipeline: a lightweight streaming module continuously processes visual input to detect changes and selectively invokes the VLM only when necessary. This decoupled design preserves the reasoning power and zero-shot generalization of VLMs while improving efficiency, avoiding both the high monetary cost and the latency of frame-by-frame captioning that leads to fragmented and delayed outputs. To address the absence of suitable benchmarks for multi-actor collaboration, we introduce the GROUND dataset, which offers fine-grained situational annotations of multi-person and human-robot interactions. On this dataset, our approach improves the average grounding score by a factor of 2 compared to the performance of VLM-only baselines - including GPT-4o, GPT-5 and Gemini 2.5 Flash - while also reducing run-time by a factor of 4. The code and data are available at www.github.com/HRI-EU/merge.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18988
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MERGE: Guided Vision-Language Models for Multi-Actor Event Reasoning and Grounding in Human-Robot Interaction
Deigmoeller, Joerg
Agarwal, Nakul
Hasler, Stephan
Tanneberg, Daniel
Belardinelli, Anna
Ghoddoosian, Reza
Wang, Chao
Ocker, Felix
Zhang, Fan
Dariush, Behzad
Gienger, Michael
Robotics
We introduce MERGE, a system for situational grounding of actors, objects, and events in dynamic human-robot group interactions. Effective collaboration in such settings requires consistent situational awareness, built on persistent representations of people and objects and an episodic abstraction of events. MERGE achieves this by uniquely identifying physical instances of actors (humans or robots) and objects and structuring them into actor-action-object relations, ensuring temporal consistency across interactions. Central to MERGE is the integration of Vision-Language Models (VLMs) guided with a perception pipeline: a lightweight streaming module continuously processes visual input to detect changes and selectively invokes the VLM only when necessary. This decoupled design preserves the reasoning power and zero-shot generalization of VLMs while improving efficiency, avoiding both the high monetary cost and the latency of frame-by-frame captioning that leads to fragmented and delayed outputs. To address the absence of suitable benchmarks for multi-actor collaboration, we introduce the GROUND dataset, which offers fine-grained situational annotations of multi-person and human-robot interactions. On this dataset, our approach improves the average grounding score by a factor of 2 compared to the performance of VLM-only baselines - including GPT-4o, GPT-5 and Gemini 2.5 Flash - while also reducing run-time by a factor of 4. The code and data are available at www.github.com/HRI-EU/merge.
title MERGE: Guided Vision-Language Models for Multi-Actor Event Reasoning and Grounding in Human-Robot Interaction
topic Robotics
url https://arxiv.org/abs/2603.18988