Affogato: Learning Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Junha, Park, Eunha, Park, Chunghyun, Kang, Dahyun, Cho, Minsu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908407381884928
author Lee, Junha
Park, Eunha
Park, Chunghyun
Kang, Dahyun
Cho, Minsu
author_facet Lee, Junha
Park, Eunha
Park, Chunghyun
Kang, Dahyun
Cho, Minsu
contents Affordance grounding-localizing object regions based on natural language descriptions of interactions-is a critical challenge for enabling intelligent agents to understand and interact with their environments. However, this task remains challenging due to the need for fine-grained part-level localization, the ambiguity arising from multiple valid interaction regions, and the scarcity of large-scale datasets. In this work, we introduce Affogato, a large-scale benchmark comprising 150K instances, annotated with open-vocabulary text descriptions and corresponding 3D affordance heatmaps across a diverse set of objects and interactions. Building on this benchmark, we develop simple yet effective vision-language models that leverage pretrained part-aware vision backbones and a text-conditional heatmap decoder. Our models trained with the Affogato dataset achieve promising performance on the existing 2D and 3D benchmarks, and notably, exhibit effectiveness in open-vocabulary cross-domain generalization. The Affogato dataset is shared in public: https://huggingface.co/datasets/project-affogato/affogato
format Preprint
id arxiv_https___arxiv_org_abs_2506_12009
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Affogato: Learning Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale
Lee, Junha
Park, Eunha
Park, Chunghyun
Kang, Dahyun
Cho, Minsu
Computer Vision and Pattern Recognition
Affordance grounding-localizing object regions based on natural language descriptions of interactions-is a critical challenge for enabling intelligent agents to understand and interact with their environments. However, this task remains challenging due to the need for fine-grained part-level localization, the ambiguity arising from multiple valid interaction regions, and the scarcity of large-scale datasets. In this work, we introduce Affogato, a large-scale benchmark comprising 150K instances, annotated with open-vocabulary text descriptions and corresponding 3D affordance heatmaps across a diverse set of objects and interactions. Building on this benchmark, we develop simple yet effective vision-language models that leverage pretrained part-aware vision backbones and a text-conditional heatmap decoder. Our models trained with the Affogato dataset achieve promising performance on the existing 2D and 3D benchmarks, and notably, exhibit effectiveness in open-vocabulary cross-domain generalization. The Affogato dataset is shared in public: https://huggingface.co/datasets/project-affogato/affogato
title Affogato: Learning Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.12009