DOCTR: Disentangled Object-Centric Transformer for Point Scene Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Xiaoxuan, Wang, Hao, Li, Weiming, Wang, Qiang, Cho, Soonyong, Sung, Younghun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Disentangled Representation in Object-Centric Models for Visual Dynamics Prediction via Transformers
von: Gandhi, Sanket, et al.
Veröffentlicht: (2024)
von: Gandhi, Sanket, et al.
Veröffentlicht: (2024)
ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives
von: Fu, Yuqian, et al.
Veröffentlicht: (2024)
von: Fu, Yuqian, et al.
Veröffentlicht: (2024)
Explicitly Disentangled Representations in Object-Centric Learning
von: Majellaro, Riccardo, et al.
Veröffentlicht: (2024)
von: Majellaro, Riccardo, et al.
Veröffentlicht: (2024)
DOR3D-Net: Dense Ordinal Regression Network for 3D Hand Pose Estimation
von: Mao, Yamin, et al.
Veröffentlicht: (2024)
von: Mao, Yamin, et al.
Veröffentlicht: (2024)
ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
von: Wang, Xucheng, et al.
Veröffentlicht: (2026)
von: Wang, Xucheng, et al.
Veröffentlicht: (2026)
TGBFormer: Transformer-GraphFormer Blender Network for Video Object Detection
von: Qi, Qiang, et al.
Veröffentlicht: (2025)
von: Qi, Qiang, et al.
Veröffentlicht: (2025)
Short-term Object Interaction Anticipation with Disentangled Object Detection @ Ego4D Short Term Object Interaction Anticipation Challenge
von: Cho, Hyunjin, et al.
Veröffentlicht: (2024)
von: Cho, Hyunjin, et al.
Veröffentlicht: (2024)
Towards Flexible 3D Perception: Object-Centric Occupancy Completion Augments 3D Object Detection
von: Zheng, Chaoda, et al.
Veröffentlicht: (2024)
von: Zheng, Chaoda, et al.
Veröffentlicht: (2024)
Object-Centric Vision Token Pruning for Vision Language Models
von: Li, Guangyuan, et al.
Veröffentlicht: (2025)
von: Li, Guangyuan, et al.
Veröffentlicht: (2025)
TripleFDS: Triple Feature Disentanglement and Synthesis for Scene Text Editing
von: Bao, Yuchen, et al.
Veröffentlicht: (2025)
von: Bao, Yuchen, et al.
Veröffentlicht: (2025)
ObjectGS: Object-aware Scene Reconstruction and Scene Understanding via Gaussian Splatting
von: Zhu, Ruijie, et al.
Veröffentlicht: (2025)
von: Zhu, Ruijie, et al.
Veröffentlicht: (2025)
OpenGS-SLAM: Open-Set Dense Semantic SLAM with 3D Gaussian Splatting for Object-Level Scene Understanding
von: Yang, Dianyi, et al.
Veröffentlicht: (2025)
von: Yang, Dianyi, et al.
Veröffentlicht: (2025)
RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding
von: Yang, Jihan, et al.
Veröffentlicht: (2023)
von: Yang, Jihan, et al.
Veröffentlicht: (2023)
InstrAct: Towards Action-Centric Understanding in Instructional Videos
von: Yang, Zhuoyi, et al.
Veröffentlicht: (2026)
von: Yang, Zhuoyi, et al.
Veröffentlicht: (2026)
UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
von: Le, Huy, et al.
Veröffentlicht: (2025)
von: Le, Huy, et al.
Veröffentlicht: (2025)
Revisiting Salient Object Detection from an Observer-Centric Perspective
von: Zhang, Fuxi, et al.
Veröffentlicht: (2026)
von: Zhang, Fuxi, et al.
Veröffentlicht: (2026)
Allocentric Perceiver: Disentangling Allocentric Reasoning from Egocentric Visual Priors via Frame Instantiation
von: Wang, Hengyi, et al.
Veröffentlicht: (2026)
von: Wang, Hengyi, et al.
Veröffentlicht: (2026)
Local Temporal Feature Enhanced Transformer with ROI-rank Based Masking for Diagnosis of ADHD
von: Kim, Byunggun, et al.
Veröffentlicht: (2025)
von: Kim, Byunggun, et al.
Veröffentlicht: (2025)
PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos
von: Zhou, Zhiyu, et al.
Veröffentlicht: (2026)
von: Zhou, Zhiyu, et al.
Veröffentlicht: (2026)
BeyondScene: Higher-Resolution Human-Centric Scene Generation With Pretrained Diffusion
von: Kim, Gwanghyun, et al.
Veröffentlicht: (2024)
von: Kim, Gwanghyun, et al.
Veröffentlicht: (2024)
Reasoning-Enhanced Object-Centric Learning for Videos
von: Li, Jian, et al.
Veröffentlicht: (2024)
von: Li, Jian, et al.
Veröffentlicht: (2024)
Dynamic Object Queries for Transformer-based Incremental Object Detection
von: Zhang, Jichuan, et al.
Veröffentlicht: (2024)
von: Zhang, Jichuan, et al.
Veröffentlicht: (2024)
Understanding Dynamic Scenes in Ego Centric 4D Point Clouds
von: Huang, Junsheng, et al.
Veröffentlicht: (2025)
von: Huang, Junsheng, et al.
Veröffentlicht: (2025)
City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning
von: Sun, Penglei, et al.
Veröffentlicht: (2025)
von: Sun, Penglei, et al.
Veröffentlicht: (2025)
Object-Centric Latent Action Learning
von: Klepach, Albina, et al.
Veröffentlicht: (2025)
von: Klepach, Albina, et al.
Veröffentlicht: (2025)
milliFlow: Scene Flow Estimation on mmWave Radar Point Cloud for Human Motion Sensing
von: Ding, Fangqiang, et al.
Veröffentlicht: (2023)
von: Ding, Fangqiang, et al.
Veröffentlicht: (2023)
Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes
von: Wang, Yaoting, et al.
Veröffentlicht: (2024)
von: Wang, Yaoting, et al.
Veröffentlicht: (2024)
CrashSight: A Phase-Aware, Infrastructure-Centric Video Benchmark for Traffic Crash Scene Understanding and Reasoning
von: Gan, Rui, et al.
Veröffentlicht: (2026)
von: Gan, Rui, et al.
Veröffentlicht: (2026)
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding
von: Zeng, Nianbo, et al.
Veröffentlicht: (2025)
von: Zeng, Nianbo, et al.
Veröffentlicht: (2025)
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
Unsupervised Anomaly Detection in Brain MRI via Disentangled Anatomy Learning
von: Yang, Tao, et al.
Veröffentlicht: (2025)
von: Yang, Tao, et al.
Veröffentlicht: (2025)
Hierarchical Object-Centric Learning with Capsule Networks
von: Renzulli, Riccardo
Veröffentlicht: (2024)
von: Renzulli, Riccardo
Veröffentlicht: (2024)
Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection
von: Seo, Soo Won, et al.
Veröffentlicht: (2026)
von: Seo, Soo Won, et al.
Veröffentlicht: (2026)
HAECcity: Open-Vocabulary Scene Understanding of City-Scale Point Clouds with Superpoint Graph Clustering
von: Rusnak, Alexander, et al.
Veröffentlicht: (2025)
von: Rusnak, Alexander, et al.
Veröffentlicht: (2025)
RPMArt: Towards Robust Perception and Manipulation for Articulated Objects
von: Wang, Junbo, et al.
Veröffentlicht: (2024)
von: Wang, Junbo, et al.
Veröffentlicht: (2024)
Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs
von: Huang, Jincai, et al.
Veröffentlicht: (2026)
von: Huang, Jincai, et al.
Veröffentlicht: (2026)
Dynamic Scene Understanding through Object-Centric Voxelization and Neural Rendering
von: Zhao, Yanpeng, et al.
Veröffentlicht: (2024)
von: Zhao, Yanpeng, et al.
Veröffentlicht: (2024)
Lifting Unlabeled Internet-level Data for 3D Scene Understanding
von: Chen, Yixin, et al.
Veröffentlicht: (2026)
von: Chen, Yixin, et al.
Veröffentlicht: (2026)
Complementary Pseudo Multimodal Feature for Point Cloud Anomaly Detection
von: Cao, Yunkang, et al.
Veröffentlicht: (2023)
von: Cao, Yunkang, et al.
Veröffentlicht: (2023)
Object-Centric 3D Gaussian Splatting for Strawberry Plant Reconstruction and Phenotyping
von: Li, Jiajia, et al.
Veröffentlicht: (2025)
von: Li, Jiajia, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Learning Disentangled Representation in Object-Centric Models for Visual Dynamics Prediction via Transformers
von: Gandhi, Sanket, et al.
Veröffentlicht: (2024) -
ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives
von: Fu, Yuqian, et al.
Veröffentlicht: (2024) -
Explicitly Disentangled Representations in Object-Centric Learning
von: Majellaro, Riccardo, et al.
Veröffentlicht: (2024) -
DOR3D-Net: Dense Ordinal Regression Network for 3D Hand Pose Estimation
von: Mao, Yamin, et al.
Veröffentlicht: (2024) -
ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
von: Wang, Xucheng, et al.
Veröffentlicht: (2026)