CRAFT: Adapting VLA Models to Contact-rich Manipulation via Force-aware Curriculum Fine-tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yike, Wang, Yaonan, Sun, Xinxin, Huang, Kaizhen, Xu, Zhiyuan, Ji, Junjie, Che, Zhengping, Tang, Jian, Sun, Jingtao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912902301089792
author Zhang, Yike
Wang, Yaonan
Sun, Xinxin
Huang, Kaizhen
Xu, Zhiyuan
Ji, Junjie
Che, Zhengping
Tang, Jian
Sun, Jingtao
author_facet Zhang, Yike
Wang, Yaonan
Sun, Xinxin
Huang, Kaizhen
Xu, Zhiyuan
Ji, Junjie
Che, Zhengping
Tang, Jian
Sun, Jingtao
contents Vision-Language-Action (VLA) models have shown a strong capability in enabling robots to execute general instructions, yet they struggle with contact-rich manipulation tasks, where success requires precise alignment, stable contact maintenance, and effective handling of deformable objects. A fundamental challenge arises from the imbalance between high-entropy vision and language inputs and low-entropy but critical force signals, which often leads to over-reliance on perception and unstable control. To address this, we introduce CRAFT, a force-aware curriculum fine-tuning framework that integrates a variational information bottleneck module to regulate vision and language embeddings during early training. This curriculum strategy encourages the model to prioritize force signals initially, before progressively restoring access to the full multimodal information. To enable force-aware learning, we further design a homologous leader-follower teleoperation system that collects synchronized vision, language, and force data across diverse contact-rich tasks. Real-world experiments demonstrate that CRAFT consistently improves task success, generalizes to unseen objects and novel task variations, and adapts effectively across diverse VLA architectures, enabling robust and generalizable contact-rich manipulation.
format Preprint
id arxiv_https___arxiv_org_abs_2602_12532
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CRAFT: Adapting VLA Models to Contact-rich Manipulation via Force-aware Curriculum Fine-tuning
Zhang, Yike
Wang, Yaonan
Sun, Xinxin
Huang, Kaizhen
Xu, Zhiyuan
Ji, Junjie
Che, Zhengping
Tang, Jian
Sun, Jingtao
Robotics
Vision-Language-Action (VLA) models have shown a strong capability in enabling robots to execute general instructions, yet they struggle with contact-rich manipulation tasks, where success requires precise alignment, stable contact maintenance, and effective handling of deformable objects. A fundamental challenge arises from the imbalance between high-entropy vision and language inputs and low-entropy but critical force signals, which often leads to over-reliance on perception and unstable control. To address this, we introduce CRAFT, a force-aware curriculum fine-tuning framework that integrates a variational information bottleneck module to regulate vision and language embeddings during early training. This curriculum strategy encourages the model to prioritize force signals initially, before progressively restoring access to the full multimodal information. To enable force-aware learning, we further design a homologous leader-follower teleoperation system that collects synchronized vision, language, and force data across diverse contact-rich tasks. Real-world experiments demonstrate that CRAFT consistently improves task success, generalizes to unseen objects and novel task variations, and adapts effectively across diverse VLA architectures, enabling robust and generalizable contact-rich manipulation.
title CRAFT: Adapting VLA Models to Contact-rich Manipulation via Force-aware Curriculum Fine-tuning
topic Robotics
url https://arxiv.org/abs/2602.12532