Ego3DT: Tracking Every 3D Object in Ego-centric Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Hao, Shengyu, Chai, Wenhao, Zhao, Zhonghan, Sun, Meiqi, Hu, Wendi, Zhou, Jieyang, Zhao, Yixian, Li, Qi, Wang, Yizhou, Li, Xi, Wang, Gaoang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unsupervised Ego- and Exo-centric Dense Procedural Activity Captioning via Gaze Consensus Adaptation
by: Shi, Zhaofeng, et al.
Published: (2025)
by: Shi, Zhaofeng, et al.
Published: (2025)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
by: Zhou, Sheng, et al.
Published: (2025)
by: Zhou, Sheng, et al.
Published: (2025)
EgoForge: Goal-Directed Egocentric World Simulator
by: Shen, Yifan, et al.
Published: (2026)
by: Shen, Yifan, et al.
Published: (2026)
EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports
by: Ma, Jianzhe, et al.
Published: (2026)
by: Ma, Jianzhe, et al.
Published: (2026)
EgoBlind: Towards Egocentric Visual Assistance for the Blind
by: Xiao, Junbin, et al.
Published: (2025)
by: Xiao, Junbin, et al.
Published: (2025)
EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos
by: Rai, Aashish, et al.
Published: (2024)
by: Rai, Aashish, et al.
Published: (2024)
SGAT4PASS: Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation
by: Li, Xuewei, et al.
Published: (2023)
by: Li, Xuewei, et al.
Published: (2023)
EgoAdapt: Enhancing Robustness in Egocentric Interactive Speaker Detection Under Missing Modalities
by: Qian, Xinyuan, et al.
Published: (2026)
by: Qian, Xinyuan, et al.
Published: (2026)
Understanding Dynamic Scenes in Ego Centric 4D Point Clouds
by: Huang, Junsheng, et al.
Published: (2025)
by: Huang, Junsheng, et al.
Published: (2025)
EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
by: Xu, Jilan, et al.
Published: (2025)
by: Xu, Jilan, et al.
Published: (2025)
MTFusion: Reconstructing Any 3D Object from Single Image Using Multi-word Textual Inversion
by: Liu, Yu, et al.
Published: (2024)
by: Liu, Yu, et al.
Published: (2024)
Startup Delay Aware Short Video Ordering: Problem, Model, and A Reinforcement Learning based Algorithm
by: Gao, Zhipeng, et al.
Published: (2024)
by: Gao, Zhipeng, et al.
Published: (2024)
PureKV: Plug-and-Play KV Cache Optimization with Spatial-Temporal Sparse Attention for Vision-Language Large Models
by: Jiang, Zhonghua, et al.
Published: (2025)
by: Jiang, Zhonghua, et al.
Published: (2025)
Priority prediction of Asian Hornet sighting report using machine learning methods
by: Liu, Yixin, et al.
Published: (2021)
by: Liu, Yixin, et al.
Published: (2021)
Every Painting Awakened: A Training-free Framework for Painting-to-Animation Generation
by: Liu, Lingyu, et al.
Published: (2025)
by: Liu, Lingyu, et al.
Published: (2025)
DreamFoley: Scalable VLMs for High-Fidelity Video-to-Audio Generation
by: Li, Fu, et al.
Published: (2025)
by: Li, Fu, et al.
Published: (2025)
Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs
by: Chen, Jianhao, et al.
Published: (2026)
by: Chen, Jianhao, et al.
Published: (2026)
Perceive-Sample-Compress: Towards Real-Time 3D Gaussian Splatting
by: Wang, Zijian, et al.
Published: (2025)
by: Wang, Zijian, et al.
Published: (2025)
Efficient Object-centric Representation Learning with Pre-trained Geometric Prior
by: Khac, Phúc H. Le, et al.
Published: (2024)
by: Khac, Phúc H. Le, et al.
Published: (2024)
Cap2Sum: Learning to Summarize Videos by Generating Captions
by: Zhao, Cairong, et al.
Published: (2024)
by: Zhao, Cairong, et al.
Published: (2024)
See and Think: Embodied Agent in Virtual Environment
by: Zhao, Zhonghan, et al.
Published: (2023)
by: Zhao, Zhonghan, et al.
Published: (2023)
Adaptive 3D Gaussian Splatting Video Streaming
by: Gong, Han, et al.
Published: (2025)
by: Gong, Han, et al.
Published: (2025)
DT-UFC: Universal Large Model Feature Coding via Peaky-to-Balanced Distribution Transformation
by: Gao, Changsheng, et al.
Published: (2025)
by: Gao, Changsheng, et al.
Published: (2025)
Casual3DHDR: Deblurring High Dynamic Range 3D Gaussian Splatting from Casually Captured Videos
by: Gong, Shucheng, et al.
Published: (2025)
by: Gong, Shucheng, et al.
Published: (2025)
Fine-grained Image Retrieval via Dual-Vision Adaptation
by: Jiang, Xin, et al.
Published: (2025)
by: Jiang, Xin, et al.
Published: (2025)
CLIPRerank: An Extremely Simple Method for Improving Ad-hoc Video Search
by: Chen, Aozhu, et al.
Published: (2024)
by: Chen, Aozhu, et al.
Published: (2024)
Music Grounding by Short Video
by: Xin, Zijie, et al.
Published: (2024)
by: Xin, Zijie, et al.
Published: (2024)
Exploring Category-level Articulated Object Pose Tracking on SE(3) Manifolds
by: Meng, Xianhui, et al.
Published: (2025)
by: Meng, Xianhui, et al.
Published: (2025)
EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations
by: Sun, Haoqin, et al.
Published: (2025)
by: Sun, Haoqin, et al.
Published: (2025)
Patch-level Sounding Object Tracking for Audio-Visual Question Answering
by: Li, Zhangbin, et al.
Published: (2024)
by: Li, Zhangbin, et al.
Published: (2024)
CDIO: Cross-Domain Inference Optimization with Resource Preference Prediction for Edge-Cloud Collaboration
by: Yang, Zheming, et al.
Published: (2025)
by: Yang, Zheming, et al.
Published: (2025)
Breaking the Curse of Knowledge: Towards Effective Multimodal Recommendation using Knowledge Soft Integration
by: Ouyang, Kai, et al.
Published: (2023)
by: Ouyang, Kai, et al.
Published: (2023)
3MDiT: Unified Tri-Modal Diffusion Transformer for Text-Driven Synchronized Audio-Video Generation
by: Li, Yaoru, et al.
Published: (2025)
by: Li, Yaoru, et al.
Published: (2025)
SpeechEE: A Novel Benchmark for Speech Event Extraction
by: Wang, Bin, et al.
Published: (2024)
by: Wang, Bin, et al.
Published: (2024)
Digital Fingerprinting on Multimedia: A Survey
by: Chen, Wendi, et al.
Published: (2024)
by: Chen, Wendi, et al.
Published: (2024)
Virbo: Multimodal Multilingual Avatar Video Generation in Digital Marketing
by: Zhang, Juan, et al.
Published: (2024)
by: Zhang, Juan, et al.
Published: (2024)
MAR3: Multi-Agent Recognition, Reasoning, and Reflection for Reference Audio-Visual Segmentation
by: Zhao, Yuan, et al.
Published: (2026)
by: Zhao, Yuan, et al.
Published: (2026)
Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline
by: Yang, Dingyi, et al.
Published: (2024)
by: Yang, Dingyi, et al.
Published: (2024)
TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing
by: Chen, Yaru, et al.
Published: (2025)
by: Chen, Yaru, et al.
Published: (2025)
T-GVC: Trajectory-Guided Generative Video Coding at Ultra-Low Bitrates
by: Wang, Zhitao, et al.
Published: (2025)
by: Wang, Zhitao, et al.
Published: (2025)
Similar Items
-
Unsupervised Ego- and Exo-centric Dense Procedural Activity Captioning via Gaze Consensus Adaptation
by: Shi, Zhaofeng, et al.
Published: (2025) -
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
by: Zhou, Sheng, et al.
Published: (2025) -
EgoForge: Goal-Directed Egocentric World Simulator
by: Shen, Yifan, et al.
Published: (2026) -
EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports
by: Ma, Jianzhe, et al.
Published: (2026) -
EgoBlind: Towards Egocentric Visual Assistance for the Blind
by: Xiao, Junbin, et al.
Published: (2025)