MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Kangsan, Yang, Yanlai, Kim, Suji, Yeo, Woongyeong, Lee, Youngwan, Ren, Mengye, Hwang, Sung Ju
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912959908806656
author Kim, Kangsan
Yang, Yanlai
Kim, Suji
Yeo, Woongyeong
Lee, Youngwan
Ren, Mengye
Hwang, Sung Ju
author_facet Kim, Kangsan
Yang, Yanlai
Kim, Suji
Yeo, Woongyeong
Lee, Youngwan
Ren, Mengye
Hwang, Sung Ju
contents As embodied models become powerful, humans will collaborate with multiple embodied AI agents at their workplace or home in the future. To ensure better communication between human users and the multi-agent system, it is crucial to interpret incoming information from agents in parallel and refer to the appropriate context for each query. Existing challenges include effectively compressing and communicating high volumes of individual sensory inputs in the form of video and correctly aggregating multiple egocentric videos to construct system-level memory. In this work, we first formally define a novel problem of understanding multiple long-horizon egocentric videos simultaneously collected from embodied agents. To facilitate research in this direction, we introduce MultiAgent-EgoQA (MA-EgoQA), a benchmark designed to systemically evaluate existing models in our scenario. MA-EgoQA provides 1.7k questions unique to multiple egocentric streams, spanning five categories: social interaction, task coordination, theory-of-mind, temporal reasoning, and environmental interaction. We further propose a simple baseline model for MA-EgoQA named EgoMAS, which leverages shared memory across embodied agents and agent-wise dynamic retrieval. Through comprehensive evaluation across diverse baselines and EgoMAS on MA-EgoQA, we find that current approaches are unable to effectively handle multiple egocentric streams, highlighting the need for future advances in system-level understanding across the agents. The code and benchmark are available at https://ma-egoqa.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2603_09827
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents
Kim, Kangsan
Yang, Yanlai
Kim, Suji
Yeo, Woongyeong
Lee, Youngwan
Ren, Mengye
Hwang, Sung Ju
Computer Vision and Pattern Recognition
Artificial Intelligence
As embodied models become powerful, humans will collaborate with multiple embodied AI agents at their workplace or home in the future. To ensure better communication between human users and the multi-agent system, it is crucial to interpret incoming information from agents in parallel and refer to the appropriate context for each query. Existing challenges include effectively compressing and communicating high volumes of individual sensory inputs in the form of video and correctly aggregating multiple egocentric videos to construct system-level memory. In this work, we first formally define a novel problem of understanding multiple long-horizon egocentric videos simultaneously collected from embodied agents. To facilitate research in this direction, we introduce MultiAgent-EgoQA (MA-EgoQA), a benchmark designed to systemically evaluate existing models in our scenario. MA-EgoQA provides 1.7k questions unique to multiple egocentric streams, spanning five categories: social interaction, task coordination, theory-of-mind, temporal reasoning, and environmental interaction. We further propose a simple baseline model for MA-EgoQA named EgoMAS, which leverages shared memory across embodied agents and agent-wise dynamic retrieval. Through comprehensive evaluation across diverse baselines and EgoMAS on MA-EgoQA, we find that current approaches are unable to effectively handle multiple egocentric streams, highlighting the need for future advances in system-level understanding across the agents. The code and benchmark are available at https://ma-egoqa.github.io.
title MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.09827