Manual-PA: Learning 3D Part Assembly from Instruction Diagrams
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jiahao, Cherian, Anoop, Rodriguez, Cristian, Deng, Weijian, Gould, Stephen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Aligning Step-by-Step Instructional Diagrams to Video Demonstrations
by: Zhang, Jiahao, et al.
Published: (2023)
by: Zhang, Jiahao, et al.
Published: (2023)
Temporally Grounding Instructional Diagrams in Unconstrained Videos
by: Zhang, Jiahao, et al.
Published: (2024)
by: Zhang, Jiahao, et al.
Published: (2024)
AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects
by: Li, Danrui, et al.
Published: (2026)
by: Li, Danrui, et al.
Published: (2026)
MMHOI: Modeling Complex 3D Multi-Human Multi-Object Interactions
by: Kogashi, Kaen, et al.
Published: (2025)
by: Kogashi, Kaen, et al.
Published: (2025)
3D-GPT: Procedural 3D Modeling with Large Language Models
by: Sun, Chunyi, et al.
Published: (2023)
by: Sun, Chunyi, et al.
Published: (2023)
Is Video Anomaly Detection Misframed? Evidence from LLM-Based and Multi-Scene Models
by: Mumcu, Furkan, et al.
Published: (2026)
by: Mumcu, Furkan, et al.
Published: (2026)
An Empirical Study Into What Matters for Calibrating Vision-Language Models
by: Tu, Weijie, et al.
Published: (2024)
by: Tu, Weijie, et al.
Published: (2024)
Leveraging Multimodal LLM Descriptions of Activity for Explainable Semi-Supervised Video Anomaly Detection
by: Mumcu, Furkan, et al.
Published: (2025)
by: Mumcu, Furkan, et al.
Published: (2025)
LLM-Guided Agentic Object Detection for Open-World Understanding
by: Mumcu, Furkan, et al.
Published: (2025)
by: Mumcu, Furkan, et al.
Published: (2025)
ComplexVAD: Detecting Interaction Anomalies in Video
by: Mumcu, Furkan, et al.
Published: (2025)
by: Mumcu, Furkan, et al.
Published: (2025)
SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic Camera
by: He, Yuhang, et al.
Published: (2024)
by: He, Yuhang, et al.
Published: (2024)
3DInAction: Understanding Human Actions in 3D Point Clouds
by: Ben-Shabat, Yizhak, et al.
Published: (2023)
by: Ben-Shabat, Yizhak, et al.
Published: (2023)
Learning-based Stage Verification System in Manual Assembly Scenarios
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
Diff3DETR:Agent-based Diffusion Model for Semi-supervised 3D Object Detection
by: Deng, Jiacheng, et al.
Published: (2024)
by: Deng, Jiacheng, et al.
Published: (2024)
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
by: Zhang, Yuyou, et al.
Published: (2025)
by: Zhang, Yuyou, et al.
Published: (2025)
Confidence and Dispersity as Signals: Unsupervised Model Evaluation and Ranking
by: Deng, Weijian, et al.
Published: (2025)
by: Deng, Weijian, et al.
Published: (2025)
Relation3D: Enhancing Relation Modeling for Point Cloud Instance Segmentation
by: Lu, Jiahao, et al.
Published: (2025)
by: Lu, Jiahao, et al.
Published: (2025)
BSNet: Box-Supervised Simulation-assisted Mean Teacher for 3D Instance Segmentation
by: Lu, Jiahao, et al.
Published: (2024)
by: Lu, Jiahao, et al.
Published: (2024)
Generative 3D Part Assembly via Part-Whole-Hierarchy Message Passing
by: Du, Bi'an, et al.
Published: (2024)
by: Du, Bi'an, et al.
Published: (2024)
CAD-Llama: Leveraging Large Language Models for Computer-Aided Design Parametric 3D Model Generation
by: Li, Jiahao, et al.
Published: (2025)
by: Li, Jiahao, et al.
Published: (2025)
Beyond the Final Layer: Hierarchical Query Fusion Transformer with Agent-Interpolation Initialization for 3D Instance Segmentation
by: Lu, Jiahao, et al.
Published: (2025)
by: Lu, Jiahao, et al.
Published: (2025)
SPAFormer: Sequential 3D Part Assembly with Transformers
by: Xu, Boshen, et al.
Published: (2024)
by: Xu, Boshen, et al.
Published: (2024)
Assembler: Scalable 3D Part Assembly via Anchor Point Diffusion
by: Zhao, Wang, et al.
Published: (2025)
by: Zhao, Wang, et al.
Published: (2025)
IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos
by: Liu, Yunong, et al.
Published: (2024)
by: Liu, Yunong, et al.
Published: (2024)
Hierarchical Part-based Generative Model for Realistic 3D Blood Vessel
by: Chen, Siqi, et al.
Published: (2025)
by: Chen, Siqi, et al.
Published: (2025)
From Instructions to Assistance: a Dataset Aligning Instruction Manuals with Assembly Videos for Evaluating Multimodal LLMs
by: Toschi, Federico, et al.
Published: (2026)
by: Toschi, Federico, et al.
Published: (2026)
LLMPhy: Parameter-Identifiable Physical Reasoning Combining Large Language Models and Physics Engines
by: Cherian, Anoop, et al.
Published: (2024)
by: Cherian, Anoop, et al.
Published: (2024)
Improving Open-World Object Localization by Discovering Background
by: Singh, Ashish, et al.
Published: (2025)
by: Singh, Ashish, et al.
Published: (2025)
Feedforward 3D Editing Learns from Semantic-Part Transformation
by: Weng, Jiawei, et al.
Published: (2026)
by: Weng, Jiawei, et al.
Published: (2026)
ChemScraper: Leveraging PDF Graphics Instructions for Molecular Diagram Parsing
by: Shah, Ayush Kumar, et al.
Published: (2023)
by: Shah, Ayush Kumar, et al.
Published: (2023)
Toward a Holistic Evaluation of Robustness in CLIP Models
by: Tu, Weijie, et al.
Published: (2024)
by: Tu, Weijie, et al.
Published: (2024)
Efficient 3D Content Reconstruction and Generation
by: Li, Jiahao
Published: (2026)
by: Li, Jiahao
Published: (2026)
SAS: Segment Any 3D Scene with Integrated 2D Priors
by: Li, Zhuoyuan, et al.
Published: (2025)
by: Li, Zhuoyuan, et al.
Published: (2025)
Dive3D: Diverse Distillation-based Text-to-3D Generation via Score Implicit Matching
by: Bai, Weimin, et al.
Published: (2025)
by: Bai, Weimin, et al.
Published: (2025)
HAL3D: Hierarchical Active Learning for Fine-Grained 3D Part Labeling
by: Yu, Fenggen, et al.
Published: (2023)
by: Yu, Fenggen, et al.
Published: (2023)
Leveraging Pretrained Diffusion Models for Zero-Shot Part Assembly
by: Zhang, Ruiyuan, et al.
Published: (2025)
by: Zhang, Ruiyuan, et al.
Published: (2025)
ARINAR: Bi-Level Autoregressive Feature-by-Feature Generative Models
by: Zhao, Qinyu, et al.
Published: (2025)
by: Zhao, Qinyu, et al.
Published: (2025)
Neuro-3D: Towards 3D Visual Decoding from EEG Signals
by: Guo, Zhanqiang, et al.
Published: (2024)
by: Guo, Zhanqiang, et al.
Published: (2024)
CRAG: Can 3D Generative Models Help 3D Assembly?
by: Jiang, Zeyu, et al.
Published: (2026)
by: Jiang, Zeyu, et al.
Published: (2026)
COPRA: Conditional Parameter Adaptation with Reinforcement Learning for Video Anomaly Detection
by: Jacob, Darryl Cherian, et al.
Published: (2026)
by: Jacob, Darryl Cherian, et al.
Published: (2026)
Similar Items
-
Aligning Step-by-Step Instructional Diagrams to Video Demonstrations
by: Zhang, Jiahao, et al.
Published: (2023) -
Temporally Grounding Instructional Diagrams in Unconstrained Videos
by: Zhang, Jiahao, et al.
Published: (2024) -
AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects
by: Li, Danrui, et al.
Published: (2026) -
MMHOI: Modeling Complex 3D Multi-Human Multi-Object Interactions
by: Kogashi, Kaen, et al.
Published: (2025) -
3D-GPT: Procedural 3D Modeling with Large Language Models
by: Sun, Chunyi, et al.
Published: (2023)