Can Transformers Capture Spatial Relations between Objects?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wen, Chuan, Jayaraman, Dinesh, Gao, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Recasting Generic Pretrained Vision Transformers As Object-Centric Scene Encoders For Manipulation Policies
von: Qian, Jianing, et al.
Veröffentlicht: (2024)
von: Qian, Jianing, et al.
Veröffentlicht: (2024)
Can We Remove the Ground? Obstacle-aware Point Cloud Compression for Remote Object Detection
von: Zeng, Pengxi, et al.
Veröffentlicht: (2024)
von: Zeng, Pengxi, et al.
Veröffentlicht: (2024)
Composing Pre-Trained Object-Centric Representations for Robotics From "What" and "Where" Foundation Models
von: Shi, Junyao, et al.
Veröffentlicht: (2024)
von: Shi, Junyao, et al.
Veröffentlicht: (2024)
SHOW3D: Capturing Scenes of 3D Hands and Objects in the Wild
von: Rim, Patrick, et al.
Veröffentlicht: (2026)
von: Rim, Patrick, et al.
Veröffentlicht: (2026)
LOC-ZSON: Language-driven Object-Centric Zero-Shot Object Retrieval and Navigation
von: Guan, Tianrui, et al.
Veröffentlicht: (2024)
von: Guan, Tianrui, et al.
Veröffentlicht: (2024)
NeuralMeshing: Complete Object Mesh Extraction from Casual Captures
von: Erich, Floris, et al.
Veröffentlicht: (2025)
von: Erich, Floris, et al.
Veröffentlicht: (2025)
Personalized Embodied Navigation for Portable Object Finding
von: Dorbala, Vishnu Sashank, et al.
Veröffentlicht: (2024)
von: Dorbala, Vishnu Sashank, et al.
Veröffentlicht: (2024)
Improving Zero-Shot ObjectNav with Generative Communication
von: Dorbala, Vishnu Sashank, et al.
Veröffentlicht: (2024)
von: Dorbala, Vishnu Sashank, et al.
Veröffentlicht: (2024)
STEP: Spatial Temporal Graph Convolutional Networks for Emotion Perception from Gaits
von: Bhattacharya, Uttaran, et al.
Veröffentlicht: (2019)
von: Bhattacharya, Uttaran, et al.
Veröffentlicht: (2019)
High-Precision Transformer-Based Visual Servoing for Humanoid Robots in Aligning Tiny Objects
von: Xue, Jialong, et al.
Veröffentlicht: (2025)
von: Xue, Jialong, et al.
Veröffentlicht: (2025)
General Flow as Foundation Affordance for Scalable Robot Learning
von: Yuan, Chengbo, et al.
Veröffentlicht: (2024)
von: Yuan, Chengbo, et al.
Veröffentlicht: (2024)
Any-point Trajectory Modeling for Policy Learning
von: Wen, Chuan, et al.
Veröffentlicht: (2023)
von: Wen, Chuan, et al.
Veröffentlicht: (2023)
DVMNet++: Rethinking Relative Pose Estimation for Unseen Objects
von: Zhao, Chen, et al.
Veröffentlicht: (2024)
von: Zhao, Chen, et al.
Veröffentlicht: (2024)
SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation
von: Shi, Hao, et al.
Veröffentlicht: (2025)
von: Shi, Hao, et al.
Veröffentlicht: (2025)
Tether: Autonomous Functional Play with Correspondence-Driven Trajectory Warping
von: Liang, William, et al.
Veröffentlicht: (2026)
von: Liang, William, et al.
Veröffentlicht: (2026)
VISO-Grasp: Vision-Language Informed Spatial Object-centric 6-DoF Active View Planning and Grasping in Clutter and Invisibility
von: Shi, Yitian, et al.
Veröffentlicht: (2025)
von: Shi, Yitian, et al.
Veröffentlicht: (2025)
Cross-Level Sensor Fusion with Object Lists via Transformer for 3D Object Detection
von: Liu, Xiangzhong, et al.
Veröffentlicht: (2025)
von: Liu, Xiangzhong, et al.
Veröffentlicht: (2025)
WD-DETR: Wavelet Denoising-Enhanced Real-Time Object Detection Transformer for Robot Perception with Event Cameras
von: Cui, Yangjie, et al.
Veröffentlicht: (2025)
von: Cui, Yangjie, et al.
Veröffentlicht: (2025)
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
von: Hanyu, Taisei, et al.
Veröffentlicht: (2025)
von: Hanyu, Taisei, et al.
Veröffentlicht: (2025)
Mash, Spread, Slice! Learning to Manipulate Object States via Visual Spatial Progress
von: Mandikal, Priyanka, et al.
Veröffentlicht: (2025)
von: Mandikal, Priyanka, et al.
Veröffentlicht: (2025)
StreamLTS: Query-based Temporal-Spatial LiDAR Fusion for Cooperative Object Detection
von: Yuan, Yunshuang, et al.
Veröffentlicht: (2024)
von: Yuan, Yunshuang, et al.
Veröffentlicht: (2024)
ZeroMimic: Distilling Robotic Manipulation Skills from Web Videos
von: Shi, Junyao, et al.
Veröffentlicht: (2025)
von: Shi, Junyao, et al.
Veröffentlicht: (2025)
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
von: Zhu, Minjie, et al.
Veröffentlicht: (2025)
von: Zhu, Minjie, et al.
Veröffentlicht: (2025)
JODA: Composable Joint Dynamics for Articulated Objects
von: Gao, Tianhong, et al.
Veröffentlicht: (2026)
von: Gao, Tianhong, et al.
Veröffentlicht: (2026)
Aligning Knowledge Graph with Visual Perception for Object-goal Navigation
von: Xu, Nuo, et al.
Veröffentlicht: (2024)
von: Xu, Nuo, et al.
Veröffentlicht: (2024)
PIRATR: Parametric Object Inference for Robotic Applications with Transformers in 3D Point Clouds
von: Schwingshackl, Michael, et al.
Veröffentlicht: (2026)
von: Schwingshackl, Michael, et al.
Veröffentlicht: (2026)
Understanding Spatio-Temporal Relations in Human-Object Interaction using Pyramid Graph Convolutional Network
von: Xing, Hao, et al.
Veröffentlicht: (2024)
von: Xing, Hao, et al.
Veröffentlicht: (2024)
Fusion-Poly: A Polyhedral Framework Based on Spatial-Temporal Fusion for 3D Multi-Object Tracking
von: Wu, Xian, et al.
Veröffentlicht: (2026)
von: Wu, Xian, et al.
Veröffentlicht: (2026)
NARF24: Estimating Articulated Object Structure for Implicit Rendering
von: Lewis, Stanley, et al.
Veröffentlicht: (2024)
von: Lewis, Stanley, et al.
Veröffentlicht: (2024)
SPLATART: Articulated Gaussian Splatting with Estimated Object Structure
von: Lewis, Stanley, et al.
Veröffentlicht: (2025)
von: Lewis, Stanley, et al.
Veröffentlicht: (2025)
Towards Precise 3D Human Pose Estimation with Multi-Perspective Spatial-Temporal Relational Transformers
von: Jiao, Jianbin, et al.
Veröffentlicht: (2024)
von: Jiao, Jianbin, et al.
Veröffentlicht: (2024)
DOFS: A Real-world 3D Deformable Object Dataset with Full Spatial Information for Dynamics Model Learning
von: Zhang, Zhen, et al.
Veröffentlicht: (2024)
von: Zhang, Zhen, et al.
Veröffentlicht: (2024)
A Monocular Event-Camera Motion Capture System
von: Bauersfeld, Leonard, et al.
Veröffentlicht: (2025)
von: Bauersfeld, Leonard, et al.
Veröffentlicht: (2025)
GMT: Goal-Conditioned Multimodal Transformer for 6-DOF Object Trajectory Synthesis in 3D Scenes
von: Zeng, Huajian, et al.
Veröffentlicht: (2026)
von: Zeng, Huajian, et al.
Veröffentlicht: (2026)
Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves
von: Zhang, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2026)
Object-Scene-Camera Decomposition and Recomposition for Data-Efficient Monocular 3D Object Detection
von: Kuang, Zhaonian, et al.
Veröffentlicht: (2026)
von: Kuang, Zhaonian, et al.
Veröffentlicht: (2026)
MotionHint: Self-Supervised Monocular Visual Odometry with Motion Constraints
von: Wang, Cong, et al.
Veröffentlicht: (2021)
von: Wang, Cong, et al.
Veröffentlicht: (2021)
Object-Centric Instruction Augmentation for Robotic Manipulation
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
JND-Guided Neural Watermarking with Spatial Transformer Decoding for Screen-Capture Robustness
von: Qin, Jiayi, et al.
Veröffentlicht: (2026)
von: Qin, Jiayi, et al.
Veröffentlicht: (2026)
Relative Position Matters: Trajectory Prediction and Planning with Polar Representation
von: Zhang, Bozhou, et al.
Veröffentlicht: (2025)
von: Zhang, Bozhou, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Recasting Generic Pretrained Vision Transformers As Object-Centric Scene Encoders For Manipulation Policies
von: Qian, Jianing, et al.
Veröffentlicht: (2024) -
Can We Remove the Ground? Obstacle-aware Point Cloud Compression for Remote Object Detection
von: Zeng, Pengxi, et al.
Veröffentlicht: (2024) -
Composing Pre-Trained Object-Centric Representations for Robotics From "What" and "Where" Foundation Models
von: Shi, Junyao, et al.
Veröffentlicht: (2024) -
SHOW3D: Capturing Scenes of 3D Hands and Objects in the Wild
von: Rim, Patrick, et al.
Veröffentlicht: (2026) -
LOC-ZSON: Language-driven Object-Centric Zero-Shot Object Retrieval and Navigation
von: Guan, Tianrui, et al.
Veröffentlicht: (2024)