MERGE: Guided Vision-Language Models for Multi-Actor Event Reasoning and Grounding in Human-Robot Interaction
Fuente:
arXiv
Saved in:
| Main Authors: | Deigmoeller, Joerg, Agarwal, Nakul, Hasler, Stephan, Tanneberg, Daniel, Belardinelli, Anna, Ghoddoosian, Reza, Wang, Chao, Ocker, Felix, Zhang, Fan, Dariush, Behzad, Gienger, Michael |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition
by: Deigmoeller, Joerg, et al.
Published: (2025)
by: Deigmoeller, Joerg, et al.
Published: (2025)
To Help or Not to Help: LLM-based Attentive Support for Human-Robot Group Interactions
by: Tanneberg, Daniel, et al.
Published: (2024)
by: Tanneberg, Daniel, et al.
Published: (2024)
LaMI: Large Language Models for Multi-Modal Human-Robot Interaction
by: Wang, Chao, et al.
Published: (2024)
by: Wang, Chao, et al.
Published: (2024)
CoPAL: Corrective Planning of Robot Actions with Large Language Models
by: Joublin, Frank, et al.
Published: (2023)
by: Joublin, Frank, et al.
Published: (2023)
Mirror Eyes: Explainable Human-Robot Interaction at a Glance
by: Krüger, Matti, et al.
Published: (2025)
by: Krüger, Matti, et al.
Published: (2025)
Pose-Aware Weakly-Supervised Action Segmentation
by: Zhao, Seth Z., et al.
Published: (2025)
by: Zhao, Seth Z., et al.
Published: (2025)
Efficient Symbolic Planning with Views
by: Hasler, Stephan, et al.
Published: (2024)
by: Hasler, Stephan, et al.
Published: (2024)
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos
by: Ghoddoosian, Reza, et al.
Published: (2024)
by: Ghoddoosian, Reza, et al.
Published: (2024)
Tulip Agent -- Enabling LLM-Based Agents to Solve Tasks Using Large Tool Libraries
by: Ocker, Felix, et al.
Published: (2024)
by: Ocker, Felix, et al.
Published: (2024)
XR$^3$: An Extended Reality Platform for Social-Physical Human-Robot Interaction
by: Wang, Chao, et al.
Published: (2026)
by: Wang, Chao, et al.
Published: (2026)
Towards Driver Behavior Understanding: Weakly-Supervised Risk Perception in Driving Scenes
by: Agarwal, Nakul, et al.
Published: (2026)
by: Agarwal, Nakul, et al.
Published: (2026)
Learning Type-Generalized Actions for Symbolic Planning
by: Tanneberg, Daniel, et al.
Published: (2023)
by: Tanneberg, Daniel, et al.
Published: (2023)
SemanticScanpath: Combining Gaze and Speech for Situated Human-Robot Interaction Using LLMs
by: Menendez, Elisabeth, et al.
Published: (2025)
by: Menendez, Elisabeth, et al.
Published: (2025)
Generation of Real-time Robotic Emotional Expressions Learning from Human Demonstration in Mixed Reality
by: Wang, Chao, et al.
Published: (2025)
by: Wang, Chao, et al.
Published: (2025)
Bio-Inspired Event-Based Visual Servoing for Ground Robots
by: Mordad, Maral, et al.
Published: (2026)
by: Mordad, Maral, et al.
Published: (2026)
Local Pairwise Distance Matching for Backpropagation-Free Reinforcement Learning
by: Tanneberg, Daniel
Published: (2025)
by: Tanneberg, Daniel
Published: (2025)
Affordance-based Robot Manipulation with Flow Matching
by: Zhang, Fan, et al.
Published: (2024)
by: Zhang, Fan, et al.
Published: (2024)
Learning Robot Manipulation from Audio World Models
by: Zhang, Fan, et al.
Published: (2025)
by: Zhang, Fan, et al.
Published: (2025)
A Grounded Memory System For Smart Personal Assistants
by: Ocker, Felix, et al.
Published: (2025)
by: Ocker, Felix, et al.
Published: (2025)
Follow the Rules: Reasoning for Video Anomaly Detection with Large Language Models
by: Yang, Yuchen, et al.
Published: (2024)
by: Yang, Yuchen, et al.
Published: (2024)
On the interference of the scattered wave and the incident wave in light scattering problems with Gaussian beams
by: Gienger, Jonas
Published: (2025)
by: Gienger, Jonas
Published: (2025)
Stacked Confusion Reject Plots (SCORE)
by: Hasler, Stephan, et al.
Published: (2024)
by: Hasler, Stephan, et al.
Published: (2024)
MotionMERGE: A Multi-granular Framework for Human Motion Editing, Reasoning, Generation, and Explanation
by: Wu, Bizhu, et al.
Published: (2026)
by: Wu, Bizhu, et al.
Published: (2026)
Interactions of Fullerene (C60) and its Hydroxyl Derivatives with Lipid Bilayer: A Coarse-Grained Molecular Dynamics Simulation
by: Dariush Mohammadyani
Published: (2014)
by: Dariush Mohammadyani
Published: (2014)
Ground States for Infrared Renormalized Translation-Invariant Non-Relativistic QED
by: Hasler, David, et al.
Published: (2022)
by: Hasler, David, et al.
Published: (2022)
Ground States for translationally invariant Pauli-Fierz Models at zero Momentum
by: Hasler, David, et al.
Published: (2020)
by: Hasler, David, et al.
Published: (2020)
Facial Emotion Learning with Text-Guided Multiview Fusion via Vision-Language Model for 3D/4D Facial Expression Recognition
by: Behzad, Muzammil
Published: (2025)
by: Behzad, Muzammil
Published: (2025)
MERGE$^3$: Efficient Evolutionary Merging on Consumer-grade GPUs
by: Mencattini, Tommaso, et al.
Published: (2025)
by: Mencattini, Tommaso, et al.
Published: (2025)
Learning Robust Reasoning through Guided Adversarial Self-Play
by: Li, Shuozhe, et al.
Published: (2026)
by: Li, Shuozhe, et al.
Published: (2026)
CCDP: Composition of Conditional Diffusion Policies with Guided Sampling
by: Razmjoo, Amirreza, et al.
Published: (2025)
by: Razmjoo, Amirreza, et al.
Published: (2025)
Lifting Vision: Ground to Aerial Localization with Reasoning Guided Planning
by: Pahari, Soham, et al.
Published: (2025)
by: Pahari, Soham, et al.
Published: (2025)
Generalized Mission Planning for Heterogeneous Multi-Robot Teams via LLM-constructed Hierarchical Trees
by: Gupta, Piyush, et al.
Published: (2025)
by: Gupta, Piyush, et al.
Published: (2025)
Reasoning-Guided Grounding: Elevating Video Anomaly Detection through Multimodal Large Language Models
by: Agarwal, Sakshi, et al.
Published: (2026)
by: Agarwal, Sakshi, et al.
Published: (2026)
MERGE: Minimal Expression-Replacement GEneralization Test for Natural Language Inference
by: Zgreabăn, Mădălina, et al.
Published: (2025)
by: Zgreabăn, Mădălina, et al.
Published: (2025)
Spatial Reasoner: A 3D Inference Pipeline for XR Applications
by: Häsler, Steven, et al.
Published: (2025)
by: Häsler, Steven, et al.
Published: (2025)
Optimal Driver Warning Generation in Dynamic Driving Environment
by: Li, Chenran, et al.
Published: (2024)
by: Li, Chenran, et al.
Published: (2024)
Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model
by: AlJunaid, Reem, et al.
Published: (2025)
by: AlJunaid, Reem, et al.
Published: (2025)
Meanings and Measurements: Multi-Agent Probabilistic Grounding for Vision-Language Navigation
by: Padhan, Swagat, et al.
Published: (2026)
by: Padhan, Swagat, et al.
Published: (2026)
Do You Need a Hand? -- a Bimanual Robotic Dressing Assistance Scheme
by: Zhu, Jihong, et al.
Published: (2023)
by: Zhu, Jihong, et al.
Published: (2023)
A Logic of General Attention Using Edge-Conditioned Event Models (Extended Version)
by: Belardinelli, Gaia, et al.
Published: (2025)
by: Belardinelli, Gaia, et al.
Published: (2025)
Similar Items
-
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition
by: Deigmoeller, Joerg, et al.
Published: (2025) -
To Help or Not to Help: LLM-based Attentive Support for Human-Robot Group Interactions
by: Tanneberg, Daniel, et al.
Published: (2024) -
LaMI: Large Language Models for Multi-Modal Human-Robot Interaction
by: Wang, Chao, et al.
Published: (2024) -
CoPAL: Corrective Planning of Robot Actions with Large Language Models
by: Joublin, Frank, et al.
Published: (2023) -
Mirror Eyes: Explainable Human-Robot Interaction at a Glance
by: Krüger, Matti, et al.
Published: (2025)