GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Dongyue, Kong, Lingdong, Huang, Tianxin, Lee, Gim Hee
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913609689333760
author Lu, Dongyue
Kong, Lingdong
Huang, Tianxin
Lee, Gim Hee
author_facet Lu, Dongyue
Kong, Lingdong
Huang, Tianxin
Lee, Gim Hee
contents Identifying affordance regions on 3D objects from semantic cues is essential for robotics and human-machine interaction. However, existing 3D affordance learning methods struggle with generalization and robustness due to limited annotated data and a reliance on 3D backbones focused on geometric encoding, which often lack resilience to real-world noise and data corruption. We propose GEAL, a novel framework designed to enhance the generalization and robustness of 3D affordance learning by leveraging large-scale pre-trained 2D models. We employ a dual-branch architecture with Gaussian splatting to establish consistent mappings between 3D point clouds and 2D representations, enabling realistic 2D renderings from sparse point clouds. A granularity-adaptive fusion module and a 2D-3D consistency alignment module further strengthen cross-modal alignment and knowledge transfer, allowing the 3D branch to benefit from the rich semantics and generalization capacity of 2D models. To holistically assess the robustness, we introduce two new corruption-based benchmarks: PIAD-C and LASO-C. Extensive experiments on public datasets and our benchmarks show that GEAL consistently outperforms existing methods across seen and novel object categories, as well as corrupted data, demonstrating robust and adaptable affordance prediction under diverse conditions. Code and corruption datasets have been made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2412_09511
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency
Lu, Dongyue
Kong, Lingdong
Huang, Tianxin
Lee, Gim Hee
Computer Vision and Pattern Recognition
Identifying affordance regions on 3D objects from semantic cues is essential for robotics and human-machine interaction. However, existing 3D affordance learning methods struggle with generalization and robustness due to limited annotated data and a reliance on 3D backbones focused on geometric encoding, which often lack resilience to real-world noise and data corruption. We propose GEAL, a novel framework designed to enhance the generalization and robustness of 3D affordance learning by leveraging large-scale pre-trained 2D models. We employ a dual-branch architecture with Gaussian splatting to establish consistent mappings between 3D point clouds and 2D representations, enabling realistic 2D renderings from sparse point clouds. A granularity-adaptive fusion module and a 2D-3D consistency alignment module further strengthen cross-modal alignment and knowledge transfer, allowing the 3D branch to benefit from the rich semantics and generalization capacity of 2D models. To holistically assess the robustness, we introduce two new corruption-based benchmarks: PIAD-C and LASO-C. Extensive experiments on public datasets and our benchmarks show that GEAL consistently outperforms existing methods across seen and novel object categories, as well as corrupted data, demonstrating robust and adaptable affordance prediction under diverse conditions. Code and corruption datasets have been made publicly available.
title GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.09511