Closed-Loop Transfer for Weakly-supervised Affordance Grounding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Jiajin, Wei, Zhengxuan, Zheng, Ge, Yang, Sibei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908603072380928
author Tang, Jiajin
Wei, Zhengxuan
Zheng, Ge
Yang, Sibei
author_facet Tang, Jiajin
Wei, Zhengxuan
Zheng, Ge
Yang, Sibei
contents Humans can perform previously unexperienced interactions with novel objects simply by observing others engage with them. Weakly-supervised affordance grounding mimics this process by learning to locate object regions that enable actions on egocentric images, using exocentric interaction images with image-level annotations. However, extracting affordance knowledge solely from exocentric images and transferring it one-way to egocentric images limits the applicability of previous works in complex interaction scenarios. Instead, this study introduces LoopTrans, a novel closed-loop framework that not only transfers knowledge from exocentric to egocentric but also transfers back to enhance exocentric knowledge extraction. Within LoopTrans, several innovative mechanisms are introduced, including unified cross-modal localization and denoising knowledge distillation, to bridge domain gaps between object-centered egocentric and interaction-centered exocentric images while enhancing knowledge transfer. Experiments show that LoopTrans achieves consistent improvements across all metrics on image and video benchmarks, even handling challenging scenarios where object interaction regions are fully occluded by the human body.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17384
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Closed-Loop Transfer for Weakly-supervised Affordance Grounding
Tang, Jiajin
Wei, Zhengxuan
Zheng, Ge
Yang, Sibei
Computer Vision and Pattern Recognition
Humans can perform previously unexperienced interactions with novel objects simply by observing others engage with them. Weakly-supervised affordance grounding mimics this process by learning to locate object regions that enable actions on egocentric images, using exocentric interaction images with image-level annotations. However, extracting affordance knowledge solely from exocentric images and transferring it one-way to egocentric images limits the applicability of previous works in complex interaction scenarios. Instead, this study introduces LoopTrans, a novel closed-loop framework that not only transfers knowledge from exocentric to egocentric but also transfers back to enhance exocentric knowledge extraction. Within LoopTrans, several innovative mechanisms are introduced, including unified cross-modal localization and denoising knowledge distillation, to bridge domain gaps between object-centered egocentric and interaction-centered exocentric images while enhancing knowledge transfer. Experiments show that LoopTrans achieves consistent improvements across all metrics on image and video benchmarks, even handling challenging scenarios where object interaction regions are fully occluded by the human body.
title Closed-Loop Transfer for Weakly-supervised Affordance Grounding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.17384