EgoAgent: A Joint Predictive Agent Model in Egocentric Worlds
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Lu, Wang, Yizhou, Tang, Shixiang, Ma, Qianhong, He, Tong, Ouyang, Wanli, Zhou, Xiaowei, Bao, Hujun, Peng, Sida |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Agent3D-Zero: An Agent for Zero-shot 3D Understanding
by: Zhang, Sha, et al.
Published: (2024)
by: Zhang, Sha, et al.
Published: (2024)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
by: Tang, Shixiang, et al.
Published: (2025)
by: Tang, Shixiang, et al.
Published: (2025)
UniVerse: Unleashing the Scene Prior of Video Diffusion Models for Robust Radiance Field Reconstruction
by: Cao, Jin, et al.
Published: (2025)
by: Cao, Jin, et al.
Published: (2025)
MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-Training
by: He, Xingyi, et al.
Published: (2025)
by: He, Xingyi, et al.
Published: (2025)
CMT: A Cascade MAR with Topology Predictor for Multimodal Conditional CAD Generation
by: Wu, Jianyu, et al.
Published: (2025)
by: Wu, Jianyu, et al.
Published: (2025)
DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM
by: Wu, Yixuan, et al.
Published: (2024)
by: Wu, Yixuan, et al.
Published: (2024)
World-Grounded Human Motion Recovery via Gravity-View Coordinates
by: Shen, Zehong, et al.
Published: (2024)
by: Shen, Zehong, et al.
Published: (2024)
Representing Long Volumetric Video with Temporal Gaussian Hierarchy
by: Xu, Zhen, et al.
Published: (2024)
by: Xu, Zhen, et al.
Published: (2024)
Ready-to-React: Online Reaction Policy for Two-Character Interaction Generation
by: Cen, Zhi, et al.
Published: (2025)
by: Cen, Zhi, et al.
Published: (2025)
Generating Human Motion in 3D Scenes from Text Descriptions
by: Cen, Zhi, et al.
Published: (2024)
by: Cen, Zhi, et al.
Published: (2024)
Multi-view Reconstruction via SfM-guided Monocular Depth Estimation
by: Guo, Haoyu, et al.
Published: (2025)
by: Guo, Haoyu, et al.
Published: (2025)
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
by: Kim, Junhyeok, et al.
Published: (2025)
by: Kim, Junhyeok, et al.
Published: (2025)
EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
by: Pei, Baoqi, et al.
Published: (2024)
by: Pei, Baoqi, et al.
Published: (2024)
Precise Action-to-Video Generation Through Visual Action Prompts
by: Wang, Yuang, et al.
Published: (2025)
by: Wang, Yuang, et al.
Published: (2025)
EgoSim: Egocentric World Simulator for Embodied Interaction Generation
by: Hao, Jinkun, et al.
Published: (2026)
by: Hao, Jinkun, et al.
Published: (2026)
Dyn-E: Local Appearance Editing of Dynamic Neural Radiance Fields
by: Zhang, Shangzan, et al.
Published: (2023)
by: Zhang, Shangzan, et al.
Published: (2023)
EgoSelf: From Memory to Personalized Egocentric Assistant
by: Wang, Yanshuo, et al.
Published: (2026)
by: Wang, Yanshuo, et al.
Published: (2026)
EgoForge: Goal-Directed Egocentric World Simulator
by: Shen, Yifan, et al.
Published: (2026)
by: Shen, Yifan, et al.
Published: (2026)
EgoLife: Towards Egocentric Life Assistant
by: Yang, Jingkang, et al.
Published: (2025)
by: Yang, Jingkang, et al.
Published: (2025)
StreamingTalker: Audio-driven 3D Facial Animation with Autoregressive Diffusion Model
by: Yang, Yifan, et al.
Published: (2025)
by: Yang, Yifan, et al.
Published: (2025)
FreeTimeGS: Free Gaussian Primitives at Anytime and Anywhere for Dynamic Scene Reconstruction
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
EnvGS: Modeling View-Dependent Appearance with Environment Gaussian
by: Xie, Tao, et al.
Published: (2024)
by: Xie, Tao, et al.
Published: (2024)
Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models
by: Jin, Yudong, et al.
Published: (2025)
by: Jin, Yudong, et al.
Published: (2025)
EgoNav: Egocentric Scene-aware Human Trajectory Prediction
by: Wang, Weizhuo, et al.
Published: (2024)
by: Wang, Weizhuo, et al.
Published: (2024)
Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation
by: Lin, Haotong, et al.
Published: (2024)
by: Lin, Haotong, et al.
Published: (2024)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
Hulk: A Universal Knowledge Translator for Human-Centric Tasks
by: Wang, Yizhou, et al.
Published: (2023)
by: Wang, Yizhou, et al.
Published: (2023)
MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents
by: Kim, Kangsan, et al.
Published: (2026)
by: Kim, Kangsan, et al.
Published: (2026)
SimpleEgo: Predicting Probabilistic Body Pose from Egocentric Cameras
by: Cuevas-Velasquez, Hanz, et al.
Published: (2024)
by: Cuevas-Velasquez, Hanz, et al.
Published: (2024)
EgoGen: An Egocentric Synthetic Data Generator
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
GVGEN: Text-to-3D Generation with Volumetric Representation
by: He, Xianglong, et al.
Published: (2024)
by: He, Xianglong, et al.
Published: (2024)
MaPa: Text-driven Photorealistic Material Painting for 3D Shapes
by: Zhang, Shangzan, et al.
Published: (2024)
by: Zhang, Shangzan, et al.
Published: (2024)
Instruct-ReID++: Towards Universal Purpose Instruction-Guided Person Re-identification
by: He, Weizhen, et al.
Published: (2024)
by: He, Weizhen, et al.
Published: (2024)
EgoM2P: Egocentric Multimodal Multitask Pretraining
by: Li, Gen, et al.
Published: (2025)
by: Li, Gen, et al.
Published: (2025)
EgoLoc: A Generalizable Solution for Temporal Interaction Localization in Egocentric Videos
by: Ma, Junyi, et al.
Published: (2025)
by: Ma, Junyi, et al.
Published: (2025)
EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World
by: Punamiya, Ryan, et al.
Published: (2026)
by: Punamiya, Ryan, et al.
Published: (2026)
EgoLCD: Egocentric Video Generation with Long Context Diffusion
by: Zhang, Liuzhou, et al.
Published: (2025)
by: Zhang, Liuzhou, et al.
Published: (2025)
Efficient LoFTR: Semi-Dense Local Feature Matching with Sparse-Like Speed
by: Wang, Yifan, et al.
Published: (2024)
by: Wang, Yifan, et al.
Published: (2024)
EgoLM: Multi-Modal Language Model of Egocentric Motions
by: Hong, Fangzhou, et al.
Published: (2024)
by: Hong, Fangzhou, et al.
Published: (2024)
EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
by: Pei, Baoqi, et al.
Published: (2025)
by: Pei, Baoqi, et al.
Published: (2025)
Similar Items
-
Agent3D-Zero: An Agent for Zero-shot 3D Understanding
by: Zhang, Sha, et al.
Published: (2024) -
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
by: Tang, Shixiang, et al.
Published: (2025) -
UniVerse: Unleashing the Scene Prior of Video Diffusion Models for Robust Radiance Field Reconstruction
by: Cao, Jin, et al.
Published: (2025) -
MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-Training
by: He, Xingyi, et al.
Published: (2025) -
CMT: A Cascade MAR with Topology Predictor for Multimodal Conditional CAD Generation
by: Wu, Jianyu, et al.
Published: (2025)