GRIM: Task-Oriented Grasping with Conditioning on Generative Examples

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shailesh, Raj, Alok, Kumar, Nayan, Shukla, Priya, Melnik, Andrew, Beetz, Michael, Nandi, Gora Chand
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911269368365056
author Shailesh
Raj, Alok
Kumar, Nayan
Shukla, Priya
Melnik, Andrew
Beetz, Michael
Nandi, Gora Chand
author_facet Shailesh
Raj, Alok
Kumar, Nayan
Shukla, Priya
Melnik, Andrew
Beetz, Michael
Nandi, Gora Chand
contents Task-Oriented Grasping (TOG) requires robots to select grasps that are functionally appropriate for a specified task - a challenge that demands an understanding of task semantics, object affordances, and functional constraints. We present GRIM (Grasp Re-alignment via Iterative Matching), a training-free framework that addresses these challenges by leveraging Video Generation Models (VGMs) together with a retrieve-align-transfer pipeline. Beyond leveraging VGMs, GRIM can construct a memory of object-task exemplars sourced from web images, human demonstrations, or generative models. The retrieved task-oriented grasp is then transferred and refined by evaluating it against a set of geometrically stable candidate grasps to ensure both functional suitability and physical feasibility. GRIM demonstrates strong generalization and achieves state-of-the-art performance on standard TOG benchmarks. Project website: https://grim-tog.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2506_15607
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GRIM: Task-Oriented Grasping with Conditioning on Generative Examples
Shailesh
Raj, Alok
Kumar, Nayan
Shukla, Priya
Melnik, Andrew
Beetz, Michael
Nandi, Gora Chand
Robotics
Task-Oriented Grasping (TOG) requires robots to select grasps that are functionally appropriate for a specified task - a challenge that demands an understanding of task semantics, object affordances, and functional constraints. We present GRIM (Grasp Re-alignment via Iterative Matching), a training-free framework that addresses these challenges by leveraging Video Generation Models (VGMs) together with a retrieve-align-transfer pipeline. Beyond leveraging VGMs, GRIM can construct a memory of object-task exemplars sourced from web images, human demonstrations, or generative models. The retrieved task-oriented grasp is then transferred and refined by evaluating it against a set of geometrically stable candidate grasps to ensure both functional suitability and physical feasibility. GRIM demonstrates strong generalization and achieves state-of-the-art performance on standard TOG benchmarks. Project website: https://grim-tog.github.io
title GRIM: Task-Oriented Grasping with Conditioning on Generative Examples
topic Robotics
url https://arxiv.org/abs/2506.15607