Manual2Skill++: Connector-Aware General Robotic Assembly from Instruction Manuals via Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Tie, Chenrui, Sun, Shengxiang, Lin, Yudi, Wang, Yanbo, Li, Zhongrui, Zhong, Zhouhan, Zhu, Jinxuan, Pang, Yiman, Chen, Haonan, Chen, Junting, Wu, Ruihai, Shao, Lin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Manual2Skill: Learning to Read Manuals and Acquire Robotic Skills for Furniture Assembly Using Vision-Language Models
by: Tie, Chenrui, et al.
Published: (2025)
by: Tie, Chenrui, et al.
Published: (2025)
AdaptPNP: Integrating Prehensile and Non-Prehensile Skills for Adaptive Robotic Manipulation
by: Zhu, Jinxuan, et al.
Published: (2025)
by: Zhu, Jinxuan, et al.
Published: (2025)
EqvAfford: SE(3) Equivariance for Point-Level Affordance Learning
by: Chen, Yue, et al.
Published: (2024)
by: Chen, Yue, et al.
Published: (2024)
ShapeForce: Low-Cost Soft Robotic Wrist for Contact-Rich Manipulation
by: Zhu, Jinxuan, et al.
Published: (2025)
by: Zhu, Jinxuan, et al.
Published: (2025)
RoTri-Diff: A Spatial Robot-Object Triadic Interaction-Guided Diffusion Model for Bimanual Manipulation
by: Chen, Zixuan, et al.
Published: (2026)
by: Chen, Zixuan, et al.
Published: (2026)
LISN: Language-Instructed Social Navigation with VLM-based Controller Modulating
by: Chen, Junting, et al.
Published: (2025)
by: Chen, Junting, et al.
Published: (2025)
AutoManual: Constructing Instruction Manuals by LLM Agents via Interactive Environmental Learning
by: Chen, Minghao, et al.
Published: (2024)
by: Chen, Minghao, et al.
Published: (2024)
Learning Part-Aware Dense 3D Feature Field for Generalizable Articulated Object Manipulation
by: Chen, Yue, et al.
Published: (2026)
by: Chen, Yue, et al.
Published: (2026)
Behavioral Cloning for Robotic Connector Assembly: An Empirical Study
by: Kernbach, Andreas, et al.
Published: (2026)
by: Kernbach, Andreas, et al.
Published: (2026)
Goal-VLA: Image-Generative VLMs as Object-Centric World Models Empowering Zero-shot Robot Manipulation
by: Chen, Haonan, et al.
Published: (2025)
by: Chen, Haonan, et al.
Published: (2025)
Learning Action Conditions from Instructional Manuals for Instruction Understanding
by: Wu, Te-Lin, et al.
Published: (2022)
by: Wu, Te-Lin, et al.
Published: (2022)
ET-SEED: Efficient Trajectory-Level SE(3) Equivariant Diffusion Policy
by: Tie, Chenrui, et al.
Published: (2024)
by: Tie, Chenrui, et al.
Published: (2024)
Manual-PA: Learning 3D Part Assembly from Instruction Diagrams
by: Zhang, Jiahao, et al.
Published: (2024)
by: Zhang, Jiahao, et al.
Published: (2024)
IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos
by: Liu, Yunong, et al.
Published: (2024)
by: Liu, Yunong, et al.
Published: (2024)
The OCLC Terminal: An Instruction Manual.
by: Arnold, Jane, et al.
Published: (1982)
by: Arnold, Jane, et al.
Published: (1982)
Instruction Manual for Catalog Production.
Published: (1970)
Published: (1970)
Learning-based Stage Verification System in Manual Assembly Scenarios
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
From Instructions to Assistance: a Dataset Aligning Instruction Manuals with Assembly Videos for Evaluating Multimodal LLMs
by: Toschi, Federico, et al.
Published: (2026)
by: Toschi, Federico, et al.
Published: (2026)
Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals
by: Wu, Te-Lin, et al.
Published: (2021)
by: Wu, Te-Lin, et al.
Published: (2021)
MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
by: Pang, Yuqi, et al.
Published: (2025)
by: Pang, Yuqi, et al.
Published: (2025)
Early Quantization Shrinks Codebook: A Simple Fix for Diversity-Preserving Tokenization
by: Zhao, Wenhao, et al.
Published: (2026)
by: Zhao, Wenhao, et al.
Published: (2026)
ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
by: Gu, Chenyang, et al.
Published: (2025)
by: Gu, Chenyang, et al.
Published: (2025)
BrickCraft: Visuomotor Skill Composition with Situated Manual Guidance for Long-Horizon Interlocking Brick Assembly
by: Yu, Jichuan, et al.
Published: (2026)
by: Yu, Jichuan, et al.
Published: (2026)
Connector-S: A Survey of Connectors in Multi-modal Large Language Models
by: Zhu, Xun, et al.
Published: (2025)
by: Zhu, Xun, et al.
Published: (2025)
Generalize by Touching: Tactile Ensemble Skill Transfer for Robotic Furniture Assembly
by: Lin, Haohong, et al.
Published: (2024)
by: Lin, Haohong, et al.
Published: (2024)
To Preserve or To Compress: An In-Depth Study of Connector Selection in Multimodal Large Language Models
by: Lin, Junyan, et al.
Published: (2024)
by: Lin, Junyan, et al.
Published: (2024)
The Block Plan. Grade Seven Instructional Manual.
by: Sigurdson, Sol E., Ed., et al.
Published: (1981)
by: Sigurdson, Sol E., Ed., et al.
Published: (1981)
Lost in Instructions: Study of Blind Users' Experiences with DIY Manuals and AI-Rewritten Instructions for Assembly, Operation, and Troubleshooting of Tangible Products
by: Reddy, Monalika Padma, et al.
Published: (2026)
by: Reddy, Monalika Padma, et al.
Published: (2026)
Bounding Shortest Closed Geodesics with Diameter on compact 2-dimensional Orbifolds Homeomorphic to $S^2$
by: Chen, Jinxuan
Published: (2024)
by: Chen, Jinxuan
Published: (2024)
Public Policy Study Skills Manual. Test Edition.
by: Coplin, William D., et al.
Published: (1985)
by: Coplin, William D., et al.
Published: (1985)
MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing
by: Zhang, Kai, et al.
Published: (2023)
by: Zhang, Kai, et al.
Published: (2023)
Internet Training Manual for Educators.
by: Chen, Li-Ling
Published: (2000)
by: Chen, Li-Ling
Published: (2000)
ManiFoundation Model for General-Purpose Robotic Manipulation of Contact Synthesis with Arbitrary Objects and Robots
by: Xu, Zhixuan, et al.
Published: (2024)
by: Xu, Zhixuan, et al.
Published: (2024)
Read Aloud Programs for the Elderly Project. Instructional Manual.
by: Leonard, Gloria, et al.
Published: (1987)
by: Leonard, Gloria, et al.
Published: (1987)
Diffraction and Scattering Aware Radio Map and Environment Reconstruction using Geometry Model-Assisted Deep Learning
by: Chen, Wangqian, et al.
Published: (2024)
by: Chen, Wangqian, et al.
Published: (2024)
AWM: Accurate Weight-Matrix Fingerprint for Large Language Models
by: Zeng, Boyi, et al.
Published: (2025)
by: Zeng, Boyi, et al.
Published: (2025)
UltraDexGrasp: Learning Universal Dexterous Grasping for Bimanual Robots with Synthetic Data
by: Yang, Sizhe, et al.
Published: (2026)
by: Yang, Sizhe, et al.
Published: (2026)
Can We Probe Spacetime Non-commutativity Through Tidal Deformability of Compact Objects?
by: Peng, Junting, et al.
Published: (2025)
by: Peng, Junting, et al.
Published: (2025)
Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs
by: Gao, Lang, et al.
Published: (2025)
by: Gao, Lang, et al.
Published: (2025)
Radio Map Assisted Approach for Interference-Aware Predictive UAV Communications
by: Li, Bowen, et al.
Published: (2024)
by: Li, Bowen, et al.
Published: (2024)
Similar Items
-
Manual2Skill: Learning to Read Manuals and Acquire Robotic Skills for Furniture Assembly Using Vision-Language Models
by: Tie, Chenrui, et al.
Published: (2025) -
AdaptPNP: Integrating Prehensile and Non-Prehensile Skills for Adaptive Robotic Manipulation
by: Zhu, Jinxuan, et al.
Published: (2025) -
EqvAfford: SE(3) Equivariance for Point-Level Affordance Learning
by: Chen, Yue, et al.
Published: (2024) -
ShapeForce: Low-Cost Soft Robotic Wrist for Contact-Rich Manipulation
by: Zhu, Jinxuan, et al.
Published: (2025) -
RoTri-Diff: A Spatial Robot-Object Triadic Interaction-Guided Diffusion Model for Bimanual Manipulation
by: Chen, Zixuan, et al.
Published: (2026)