HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bao, Chen, Xu, Jiarui, Wang, Xiaolong, Gupta, Abhinav, Bharadhwaj, Homanga
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929637410471936
author Bao, Chen
Xu, Jiarui
Wang, Xiaolong
Gupta, Abhinav
Bharadhwaj, Homanga
author_facet Bao, Chen
Xu, Jiarui
Wang, Xiaolong
Gupta, Abhinav
Bharadhwaj, Homanga
contents How can we predict future interaction trajectories of human hands in a scene given high-level colloquial task specifications in the form of natural language? In this paper, we extend the classic hand trajectory prediction task to two tasks involving explicit or implicit language queries. Our proposed tasks require extensive understanding of human daily activities and reasoning abilities about what should be happening next given cues from the current scene. We also develop new benchmarks to evaluate the proposed two tasks, Vanilla Hand Prediction (VHP) and Reasoning-Based Hand Prediction (RBHP). We enable solving these tasks by integrating high-level world knowledge and reasoning capabilities of Vision-Language Models (VLMs) with the auto-regressive nature of low-level ego-centric hand trajectories. Our model, HandsOnVLM is a novel VLM that can generate textual responses and produce future hand trajectories through natural-language conversations. Our experiments show that HandsOnVLM outperforms existing task-specific methods and other VLM baselines on proposed tasks, and demonstrates its ability to effectively utilize world knowledge for reasoning about low-level human hand trajectories based on the provided context. Our website contains code and detailed video results https://www.chenbao.tech/handsonvlm/
format Preprint
id arxiv_https___arxiv_org_abs_2412_13187
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
Bao, Chen
Xu, Jiarui
Wang, Xiaolong
Gupta, Abhinav
Bharadhwaj, Homanga
Computer Vision and Pattern Recognition
Machine Learning
How can we predict future interaction trajectories of human hands in a scene given high-level colloquial task specifications in the form of natural language? In this paper, we extend the classic hand trajectory prediction task to two tasks involving explicit or implicit language queries. Our proposed tasks require extensive understanding of human daily activities and reasoning abilities about what should be happening next given cues from the current scene. We also develop new benchmarks to evaluate the proposed two tasks, Vanilla Hand Prediction (VHP) and Reasoning-Based Hand Prediction (RBHP). We enable solving these tasks by integrating high-level world knowledge and reasoning capabilities of Vision-Language Models (VLMs) with the auto-regressive nature of low-level ego-centric hand trajectories. Our model, HandsOnVLM is a novel VLM that can generate textual responses and produce future hand trajectories through natural-language conversations. Our experiments show that HandsOnVLM outperforms existing task-specific methods and other VLM baselines on proposed tasks, and demonstrates its ability to effectively utilize world knowledge for reasoning about low-level human hand trajectories based on the provided context. Our website contains code and detailed video results https://www.chenbao.tech/handsonvlm/
title HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2412.13187