Task-Aware 3D Affordance Segmentation via 2D Guidance and Geometric Refinement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Lian, Liu, Meng, Ye, Qilang, Zhou, Yu, Deng, Xiang, Ding, Gangyi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912710252298240
author He, Lian
Liu, Meng
Ye, Qilang
Zhou, Yu
Deng, Xiang
Ding, Gangyi
author_facet He, Lian
Liu, Meng
Ye, Qilang
Zhou, Yu
Deng, Xiang
Ding, Gangyi
contents Understanding 3D scene-level affordances from natural language instructions is essential for enabling embodied agents to interact meaningfully in complex environments. However, this task remains challenging due to the need for semantic reasoning and spatial grounding. Existing methods mainly focus on object-level affordances or merely lift 2D predictions to 3D, neglecting rich geometric structure information in point clouds and incurring high computational costs. To address these limitations, we introduce Task-Aware 3D Scene-level Affordance segmentation (TASA), a novel geometry-optimized framework that jointly leverages 2D semantic cues and 3D geometric reasoning in a coarse-to-fine manner. To improve the affordance detection efficiency, TASA features a task-aware 2D affordance detection module to identify manipulable points from language and visual inputs, guiding the selection of task-relevant views. To fully exploit 3D geometric information, a 3D affordance refinement module is proposed to integrate 2D semantic priors with local 3D geometry, resulting in accurate and spatially coherent 3D affordance masks. Experiments on SceneFun3D demonstrate that TASA significantly outperforms the baselines in both accuracy and efficiency in scene-level affordance segmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2511_11702
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Task-Aware 3D Affordance Segmentation via 2D Guidance and Geometric Refinement
He, Lian
Liu, Meng
Ye, Qilang
Zhou, Yu
Deng, Xiang
Ding, Gangyi
Computer Vision and Pattern Recognition
Artificial Intelligence
Image and Video Processing
Understanding 3D scene-level affordances from natural language instructions is essential for enabling embodied agents to interact meaningfully in complex environments. However, this task remains challenging due to the need for semantic reasoning and spatial grounding. Existing methods mainly focus on object-level affordances or merely lift 2D predictions to 3D, neglecting rich geometric structure information in point clouds and incurring high computational costs. To address these limitations, we introduce Task-Aware 3D Scene-level Affordance segmentation (TASA), a novel geometry-optimized framework that jointly leverages 2D semantic cues and 3D geometric reasoning in a coarse-to-fine manner. To improve the affordance detection efficiency, TASA features a task-aware 2D affordance detection module to identify manipulable points from language and visual inputs, guiding the selection of task-relevant views. To fully exploit 3D geometric information, a 3D affordance refinement module is proposed to integrate 2D semantic priors with local 3D geometry, resulting in accurate and spatially coherent 3D affordance masks. Experiments on SceneFun3D demonstrate that TASA significantly outperforms the baselines in both accuracy and efficiency in scene-level affordance segmentation.
title Task-Aware 3D Affordance Segmentation via 2D Guidance and Geometric Refinement
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Image and Video Processing
url https://arxiv.org/abs/2511.11702