SPAFormer: Sequential 3D Part Assembly with Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Boshen, Zheng, Sipeng, Jin, Qin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
POV: Prompt-Oriented View-Agnostic Learning for Egocentric Hand-Object Interaction in the Multi-View World
by: Xu, Boshen, et al.
Published: (2024)
by: Xu, Boshen, et al.
Published: (2024)
EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining
by: Xu, Boshen, et al.
Published: (2025)
by: Xu, Boshen, et al.
Published: (2025)
Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?
by: Xu, Boshen, et al.
Published: (2024)
by: Xu, Boshen, et al.
Published: (2024)
LiDAR-EVS: Enhance Extrapolated View Synthesis for 3D Gaussian Splatting with Pseudo-LiDAR Supervision
by: Huang, Yiming, et al.
Published: (2026)
by: Huang, Yiming, et al.
Published: (2026)
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
by: Luo, Hao, et al.
Published: (2025)
by: Luo, Hao, et al.
Published: (2025)
MUT3R: Motion-aware Updating Transformer for Dynamic 3D Reconstruction
by: Shen, Guole, et al.
Published: (2025)
by: Shen, Guole, et al.
Published: (2025)
AdaOcc: Adaptive Forward View Transformation and Flow Modeling for 3D Occupancy and Flow Prediction
by: Chen, Dubing, et al.
Published: (2024)
by: Chen, Dubing, et al.
Published: (2024)
Part-Guided 3D RL for Sim2Real Articulated Object Manipulation
by: Xie, Pengwei, et al.
Published: (2024)
by: Xie, Pengwei, et al.
Published: (2024)
Structural Action Transformer for 3D Dexterous Manipulation
by: Lei, Xiaohan, et al.
Published: (2026)
by: Lei, Xiaohan, et al.
Published: (2026)
ToosiCubix: Monocular 3D Cuboid Labeling via Vehicle Part Annotations
by: Nasihatkon, Behrooz, et al.
Published: (2025)
by: Nasihatkon, Behrooz, et al.
Published: (2025)
A Generalization of CLAP from 3D Localization to Image Processing, A Connection With RANSAC & Hough Transforms
by: Hou, Ruochen, et al.
Published: (2025)
by: Hou, Ruochen, et al.
Published: (2025)
Multi-robot autonomous 3D reconstruction using Gaussian splatting with Semantic guidance
by: Zeng, Jing, et al.
Published: (2024)
by: Zeng, Jing, et al.
Published: (2024)
Being-H0.7: A Latent World-Action Model from Egocentric Videos
by: Luo, Hao, et al.
Published: (2026)
by: Luo, Hao, et al.
Published: (2026)
MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning
by: Hao, Jinkun, et al.
Published: (2025)
by: Hao, Jinkun, et al.
Published: (2025)
Transforming Omnidirectional RGB-LiDAR data into 3D Gaussian Splatting
by: Bae, Semin, et al.
Published: (2026)
by: Bae, Semin, et al.
Published: (2026)
OmniIndoor3D: Comprehensive Indoor 3D Reconstruction
by: Wei, Xiaobao, et al.
Published: (2025)
by: Wei, Xiaobao, et al.
Published: (2025)
Enhancing Scene Coordinate Regression with Efficient Keypoint Detection and Sequential Information
by: Xu, Kuan, et al.
Published: (2024)
by: Xu, Kuan, et al.
Published: (2024)
Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot Learning
by: Liang, Qiwei, et al.
Published: (2025)
by: Liang, Qiwei, et al.
Published: (2025)
CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object Detection
by: Kuang, Zhaonian, et al.
Published: (2026)
by: Kuang, Zhaonian, et al.
Published: (2026)
PIRATR: Parametric Object Inference for Robotic Applications with Transformers in 3D Point Clouds
by: Schwingshackl, Michael, et al.
Published: (2026)
by: Schwingshackl, Michael, et al.
Published: (2026)
ASDF: Assembly State Detection Utilizing Late Fusion by Integrating 6D Pose Estimation
by: Schieber, Hannah, et al.
Published: (2024)
by: Schieber, Hannah, et al.
Published: (2024)
R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation
by: Xu, Xiuwei, et al.
Published: (2025)
by: Xu, Xiuwei, et al.
Published: (2025)
R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation
by: Zhang, Yuhao, et al.
Published: (2026)
by: Zhang, Yuhao, et al.
Published: (2026)
Theoretical Analysis for Expectation-Maximization-Based Multi-Model 3D Registration
by: Jin, David, et al.
Published: (2024)
by: Jin, David, et al.
Published: (2024)
GaussianGrasper: 3D Language Gaussian Splatting for Open-vocabulary Robotic Grasping
by: Zheng, Yuhang, et al.
Published: (2024)
by: Zheng, Yuhang, et al.
Published: (2024)
Cross-Level Sensor Fusion with Object Lists via Transformer for 3D Object Detection
by: Liu, Xiangzhong, et al.
Published: (2025)
by: Liu, Xiangzhong, et al.
Published: (2025)
Fish2Mesh Transformer: 3D Human Mesh Recovery from Egocentric Vision
by: Jeong, David C., et al.
Published: (2025)
by: Jeong, David C., et al.
Published: (2025)
Automating Deformable Gasket Assembly
by: Adebola, Simeon, et al.
Published: (2024)
by: Adebola, Simeon, et al.
Published: (2024)
Tele-Catch: Adaptive Teleoperation for Dexterous Dynamic 3D Object Catching
by: Zhao, Weiguang, et al.
Published: (2026)
by: Zhao, Weiguang, et al.
Published: (2026)
GMT: Goal-Conditioned Multimodal Transformer for 6-DOF Object Trajectory Synthesis in 3D Scenes
by: Zeng, Huajian, et al.
Published: (2026)
by: Zeng, Huajian, et al.
Published: (2026)
Learning Part-Aware Dense 3D Feature Field for Generalizable Articulated Object Manipulation
by: Chen, Yue, et al.
Published: (2026)
by: Chen, Yue, et al.
Published: (2026)
MSSF: A 4D Radar and Camera Fusion Framework With Multi-Stage Sampling for 3D Object Detection in Autonomous Driving
by: Liu, Hongsi, et al.
Published: (2024)
by: Liu, Hongsi, et al.
Published: (2024)
MonoDiff9D: Monocular Category-Level 9D Object Pose Estimation via Diffusion Model
by: Liu, Jian, et al.
Published: (2025)
by: Liu, Jian, et al.
Published: (2025)
CO^3: Cooperative Unsupervised 3D Representation Learning for Autonomous Driving
by: Chen, Runjian, et al.
Published: (2022)
by: Chen, Runjian, et al.
Published: (2022)
Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot Manipulation
by: Qi, Yu, et al.
Published: (2025)
by: Qi, Yu, et al.
Published: (2025)
History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation
by: Ding, Xichen, et al.
Published: (2025)
by: Ding, Xichen, et al.
Published: (2025)
Component Selection for Craft Assembly Tasks
by: Isume, Vitor Hideyo, et al.
Published: (2024)
by: Isume, Vitor Hideyo, et al.
Published: (2024)
Dynamic Visual SLAM using a General 3D Prior
by: Zhong, Xingguang, et al.
Published: (2025)
by: Zhong, Xingguang, et al.
Published: (2025)
SHOW3D: Capturing Scenes of 3D Hands and Objects in the Wild
by: Rim, Patrick, et al.
Published: (2026)
by: Rim, Patrick, et al.
Published: (2026)
NLiPsCalib: An Efficient Calibration Framework for High-Fidelity 3D Reconstruction of Curved Visuotactile Sensors
by: Qin, Xuhao, et al.
Published: (2026)
by: Qin, Xuhao, et al.
Published: (2026)
Similar Items
-
POV: Prompt-Oriented View-Agnostic Learning for Egocentric Hand-Object Interaction in the Multi-View World
by: Xu, Boshen, et al.
Published: (2024) -
EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining
by: Xu, Boshen, et al.
Published: (2025) -
Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?
by: Xu, Boshen, et al.
Published: (2024) -
LiDAR-EVS: Enhance Extrapolated View Synthesis for 3D Gaussian Splatting with Pseudo-LiDAR Supervision
by: Huang, Yiming, et al.
Published: (2026) -
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
by: Luo, Hao, et al.
Published: (2025)