DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yani, Wu, Dongming, Shi, Hao, Liu, Yingfei, Wang, Tiancai, Dong, Xingping |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bootstrapping Referring Multi-Object Tracking
by: Zhang, Yani, et al.
Published: (2024)
by: Zhang, Yani, et al.
Published: (2024)
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding
by: Zheng, Henry, et al.
Published: (2025)
by: Zheng, Henry, et al.
Published: (2025)
Ego3DT: Tracking Every 3D Object in Ego-centric Videos
by: Hao, Shengyu, et al.
Published: (2024)
by: Hao, Shengyu, et al.
Published: (2024)
Language Prompt for Autonomous Driving
by: Wu, Dongming, et al.
Published: (2023)
by: Wu, Dongming, et al.
Published: (2023)
Is a 3D-Tokenized LLM the Key to Reliable Autonomous Driving?
by: Bai, Yifan, et al.
Published: (2024)
by: Bai, Yifan, et al.
Published: (2024)
EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
by: Xu, Jilan, et al.
Published: (2025)
by: Xu, Jilan, et al.
Published: (2025)
From Web to Pixels: Bringing Agentic Search into Visual Perception
by: Yang, Bokang, et al.
Published: (2026)
by: Yang, Bokang, et al.
Published: (2026)
Ego-centric Predictive Model Conditioned on Hand Trajectories
by: Zhang, Binjie, et al.
Published: (2025)
by: Zhang, Binjie, et al.
Published: (2025)
EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World
by: Huang, Yifei, et al.
Published: (2024)
by: Huang, Yifei, et al.
Published: (2024)
Towards Visual Query Localization in the 3D World
by: Peng, Liang, et al.
Published: (2026)
by: Peng, Liang, et al.
Published: (2026)
A Simple and Better Baseline for Visual Grounding
by: Wang, Jingchao, et al.
Published: (2025)
by: Wang, Jingchao, et al.
Published: (2025)
SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation
by: Shi, Hao, et al.
Published: (2025)
by: Shi, Hao, et al.
Published: (2025)
EgoPoseFormer: A Simple Baseline for Stereo Egocentric 3D Human Pose Estimation
by: Yang, Chenhongyi, et al.
Published: (2024)
by: Yang, Chenhongyi, et al.
Published: (2024)
Object-centric Video Question Answering with Visual Grounding and Referring
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models
by: Cao, Meng, et al.
Published: (2024)
by: Cao, Meng, et al.
Published: (2024)
RoboMatrix: A Skill-centric Hierarchical Framework for Scalable Robot Task Planning and Execution in Open-World
by: Mao, Weixin, et al.
Published: (2024)
by: Mao, Weixin, et al.
Published: (2024)
Test-time Ego-Exo-centric Adaptation for Action Anticipation via Multi-Label Prototype Growing and Dual-Clue Consistency
by: Shi, Zhaofeng, et al.
Published: (2026)
by: Shi, Zhaofeng, et al.
Published: (2026)
An Efficient and Effective Transformer Decoder-Based Framework for Multi-Task Visual Grounding
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
Joint Top-Down and Bottom-Up Frameworks for 3D Visual Grounding
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment
by: Li, Xiaoqi, et al.
Published: (2025)
by: Li, Xiaoqi, et al.
Published: (2025)
RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping
by: Wu, Dongming, et al.
Published: (2025)
by: Wu, Dongming, et al.
Published: (2025)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
by: Shi, Hao, et al.
Published: (2025)
by: Shi, Hao, et al.
Published: (2025)
EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning
by: Xie, Binzhu, et al.
Published: (2026)
by: Xie, Binzhu, et al.
Published: (2026)
UniScene: Unified Occupancy-centric Driving Scene Generation
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving
by: Li, Fuhao, et al.
Published: (2025)
by: Li, Fuhao, et al.
Published: (2025)
Stream Query Denoising for Vectorized HD Map Construction
by: Wang, Shuo, et al.
Published: (2024)
by: Wang, Shuo, et al.
Published: (2024)
EgoAVU: Egocentric Audio-Visual Understanding
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
EgoSplat: Open-Vocabulary Egocentric Scene Understanding with Language Embedded 3D Gaussian Splatting
by: Li, Di, et al.
Published: (2025)
by: Li, Di, et al.
Published: (2025)
PD-APE: A Parallel Decoding Framework with Adaptive Position Encoding for 3D Visual Grounding
by: Hou, Chenshu, et al.
Published: (2024)
by: Hou, Chenshu, et al.
Published: (2024)
HCQA @ Ego4D EgoSchema Challenge 2024
by: Zhang, Haoyu, et al.
Published: (2024)
by: Zhang, Haoyu, et al.
Published: (2024)
Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
by: Zheng, Anlin, et al.
Published: (2025)
by: Zheng, Anlin, et al.
Published: (2025)
Towards High-Fidelity 3D Portrait Generation with Rich Details by Cross-View Prior-Aware Diffusion
by: Wei, Haoran, et al.
Published: (2024)
by: Wei, Haoran, et al.
Published: (2024)
ChangingGrounding: 3D Visual Grounding in Changing Scenes
by: Hu, Miao, et al.
Published: (2025)
by: Hu, Miao, et al.
Published: (2025)
The Invisible EgoHand: 3D Hand Forecasting through EgoBody Pose Estimation
by: Hatano, Masashi, et al.
Published: (2025)
by: Hatano, Masashi, et al.
Published: (2025)
Human-AI Divergence in Ego-centric Action Recognition under Spatial and Spatiotemporal Manipulations
by: Rahmaniboldaji, Sadegh, et al.
Published: (2026)
by: Rahmaniboldaji, Sadegh, et al.
Published: (2026)
PADriver: Towards Personalized Autonomous Driving
by: Kou, Genghua, et al.
Published: (2025)
by: Kou, Genghua, et al.
Published: (2025)
Panacea+: Panoramic and Controllable Video Generation for Autonomous Driving
by: Wen, Yuqing, et al.
Published: (2024)
by: Wen, Yuqing, et al.
Published: (2024)
Motor Focus: Fast Ego-Motion Prediction for Assistive Visual Navigation
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
MAC-Ego3D: Multi-Agent Gaussian Consensus for Real-Time Collaborative Ego-Motion and Photorealistic 3D Reconstruction
by: Xu, Xiaohao, et al.
Published: (2024)
by: Xu, Xiaohao, et al.
Published: (2024)
Similar Items
-
Bootstrapping Referring Multi-Object Tracking
by: Zhang, Yani, et al.
Published: (2024) -
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding
by: Zheng, Henry, et al.
Published: (2025) -
Ego3DT: Tracking Every 3D Object in Ego-centric Videos
by: Hao, Shengyu, et al.
Published: (2024) -
Language Prompt for Autonomous Driving
by: Wu, Dongming, et al.
Published: (2023) -
Is a 3D-Tokenized LLM the Key to Reliable Autonomous Driving?
by: Bai, Yifan, et al.
Published: (2024)