EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pei, Baoqi, Huang, Yifei, Xu, Jilan, He, Yuping, Chen, Guo, Wu, Fei, Qiao, Yu, Pang, Jiangmiao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912672745783296
author Pei, Baoqi
Huang, Yifei
Xu, Jilan
He, Yuping
Chen, Guo
Wu, Fei
Qiao, Yu
Pang, Jiangmiao
author_facet Pei, Baoqi
Huang, Yifei
Xu, Jilan
He, Yuping
Chen, Guo
Wu, Fei
Qiao, Yu
Pang, Jiangmiao
contents Egocentric video reasoning centers on an unobservable agent behind the camera who dynamically shapes the environment, requiring inference of hidden intentions and recognition of fine-grained interactions. This core challenge limits current multimodal large language models MLLMs, which excel at visible event reasoning but lack embodied, first-person understanding. To bridge this gap, we introduce EgoThinker, a novel framework that endows MLLMs with robust egocentric reasoning capabilities through spatio-temporal chain-of-thought supervision and a two-stage learning curriculum. First, we introduce EgoRe-5M, a large-scale egocentric QA dataset constructed from 13M diverse egocentric video clips. This dataset features multi-minute segments annotated with detailed CoT rationales and dense hand-object grounding. Second, we employ SFT on EgoRe-5M to instill reasoning skills, followed by reinforcement fine-tuning RFT to further enhance spatio-temporal localization. Experimental results show that EgoThinker outperforms existing methods across multiple egocentric benchmarks, while achieving substantial improvements in fine-grained spatio-temporal localization tasks. Full code and data are released at https://github.com/InternRobotics/EgoThinker.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23569
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
Pei, Baoqi
Huang, Yifei
Xu, Jilan
He, Yuping
Chen, Guo
Wu, Fei
Qiao, Yu
Pang, Jiangmiao
Computer Vision and Pattern Recognition
Egocentric video reasoning centers on an unobservable agent behind the camera who dynamically shapes the environment, requiring inference of hidden intentions and recognition of fine-grained interactions. This core challenge limits current multimodal large language models MLLMs, which excel at visible event reasoning but lack embodied, first-person understanding. To bridge this gap, we introduce EgoThinker, a novel framework that endows MLLMs with robust egocentric reasoning capabilities through spatio-temporal chain-of-thought supervision and a two-stage learning curriculum. First, we introduce EgoRe-5M, a large-scale egocentric QA dataset constructed from 13M diverse egocentric video clips. This dataset features multi-minute segments annotated with detailed CoT rationales and dense hand-object grounding. Second, we employ SFT on EgoRe-5M to instill reasoning skills, followed by reinforcement fine-tuning RFT to further enhance spatio-temporal localization. Experimental results show that EgoThinker outperforms existing methods across multiple egocentric benchmarks, while achieving substantial improvements in fine-grained spatio-temporal localization tasks. Full code and data are released at https://github.com/InternRobotics/EgoThinker.
title EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.23569