Object Affordance Recognition and Grounding via Multi-scale Cross-modal Representation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wan, Xinhang, Gou, Dongqiang, Liu, Xinwang, Zhu, En, He, Xuming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Part-Aware Open-Vocabulary 3D Affordance Grounding via Prototypical Semantic and Geometric Alignment
by: Gou, Dongqiang, et al.
Published: (2026)
by: Gou, Dongqiang, et al.
Published: (2026)
Intra-view and Inter-view Correlation Guided Multi-view Novel Class Discovery
by: Wan, Xinhang, et al.
Published: (2025)
by: Wan, Xinhang, et al.
Published: (2025)
Generalized Deep Multi-view Clustering via Causal Learning with Partially Aligned Cross-view Correspondence
by: Yang, Xihong, et al.
Published: (2025)
by: Yang, Xihong, et al.
Published: (2025)
Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion Recognition
by: He, Wen-Jue, et al.
Published: (2025)
by: He, Wen-Jue, et al.
Published: (2025)
Contrastive Continual Multi-view Clustering with Filtered Structural Fusion
by: Wan, Xinhang, et al.
Published: (2023)
by: Wan, Xinhang, et al.
Published: (2023)
Spectral Discrepancy and Cross-modal Semantic Consistency Learning for Object Detection in Hyperspectral Image
by: He, Xiao, et al.
Published: (2025)
by: He, Xiao, et al.
Published: (2025)
AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesis
by: Wu, Xiaofei, et al.
Published: (2026)
by: Wu, Xiaofei, et al.
Published: (2026)
VAGNet: Grounding 3D Affordance from Human-Object Interactions in Videos
by: Mao, Aihua, et al.
Published: (2026)
by: Mao, Aihua, et al.
Published: (2026)
Learning Disentangled Representations for Generalized Multi-view Clustering
by: Zou, Xin, et al.
Published: (2026)
by: Zou, Xin, et al.
Published: (2026)
MAG-VLAQ: Multi-modal Aerial-Ground Query Aggregation for Cross-View Place Recognition
by: Xu, Zhengyi, et al.
Published: (2026)
by: Xu, Zhengyi, et al.
Published: (2026)
CrossTracker: Robust Multi-modal 3D Multi-Object Tracking via Cross Correction
by: Gu, Lipeng, et al.
Published: (2024)
by: Gu, Lipeng, et al.
Published: (2024)
Grounding by Remembering: Cross-Scene and In-Scene Memory for 3D Functional Affordances
by: Wang, Qirui, et al.
Published: (2026)
by: Wang, Qirui, et al.
Published: (2026)
Incremental Human-Object Interaction Detection with Invariant Relation Representation Learning
by: Wei, Yana, et al.
Published: (2025)
by: Wei, Yana, et al.
Published: (2025)
LVLM-empowered Multi-modal Representation Learning for Visual Place Recognition
by: Wang, Teng, et al.
Published: (2024)
by: Wang, Teng, et al.
Published: (2024)
Affordance-Aware Object Insertion via Mask-Aware Dual Diffusion
by: He, Jixuan, et al.
Published: (2024)
by: He, Jixuan, et al.
Published: (2024)
Learning 2D Invariant Affordance Knowledge for 3D Affordance Grounding
by: Gao, Xianqiang, et al.
Published: (2024)
by: Gao, Xianqiang, et al.
Published: (2024)
DAP: Diffusion-based Affordance Prediction for Multi-modality Storage
by: Chang, Haonan, et al.
Published: (2024)
by: Chang, Haonan, et al.
Published: (2024)
Learning from Observer Gaze:Zero-Shot Attention Prediction Oriented by Human-Object Interaction Recognition
by: Zhou, Yuchen, et al.
Published: (2024)
by: Zhou, Yuchen, et al.
Published: (2024)
HAMMER: Harnessing MLLM via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding
by: Yao, Lei, et al.
Published: (2026)
by: Yao, Lei, et al.
Published: (2026)
Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
by: Zhu, He, et al.
Published: (2025)
by: Zhu, He, et al.
Published: (2025)
AffordanceSAM: Segment Anything Once More in Affordance Grounding
by: Jiang, Dengyang, et al.
Published: (2025)
by: Jiang, Dengyang, et al.
Published: (2025)
Phrase Grounding-based Style Transfer for Single-Domain Generalized Object Detection
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model
by: Wang, Hanqing, et al.
Published: (2026)
by: Wang, Hanqing, et al.
Published: (2026)
Cross-domain Multi-modal Few-shot Object Detection via Rich Text
by: Shangguan, Zeyu, et al.
Published: (2024)
by: Shangguan, Zeyu, et al.
Published: (2024)
Deep Incomplete Multi-view Clustering with Distribution Dual-Consistency Recovery Guidance
by: Jin, Jiaqi, et al.
Published: (2025)
by: Jin, Jiaqi, et al.
Published: (2025)
Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement
by: Li, Bing, et al.
Published: (2022)
by: Li, Bing, et al.
Published: (2022)
Improving Out-of-distribution Human Activity Recognition via IMU-Video Cross-modal Representation Learning
by: Cheshmi, Seyyed Saeid, et al.
Published: (2025)
by: Cheshmi, Seyyed Saeid, et al.
Published: (2025)
Multi-modal Generation via Cross-Modal In-Context Learning
by: Kumar, Amandeep, et al.
Published: (2024)
by: Kumar, Amandeep, et al.
Published: (2024)
AffordanceLLM: Grounding Affordance from Vision Language Models
by: Qian, Shengyi, et al.
Published: (2024)
by: Qian, Shengyi, et al.
Published: (2024)
View-aware Cross-modal Distillation for Multi-view Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2025)
by: Nguyen, Trung Thanh, et al.
Published: (2025)
Robust Domain Generalization for Multi-modal Object Recognition
by: Qiao, Yuxin, et al.
Published: (2024)
by: Qiao, Yuxin, et al.
Published: (2024)
Cross-View Referring Multi-Object Tracking
by: Chen, Sijia, et al.
Published: (2024)
by: Chen, Sijia, et al.
Published: (2024)
Learning Visual Affordance from Audio
by: Lu, Lidong, et al.
Published: (2025)
by: Lu, Lidong, et al.
Published: (2025)
EventDance: Unsupervised Source-free Cross-modal Adaptation for Event-based Object Recognition
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
Automatically Identify and Rectify: Robust Deep Contrastive Multi-view Clustering in Noisy Scenarios
by: Yang, Xihong, et al.
Published: (2025)
by: Yang, Xihong, et al.
Published: (2025)
CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects
by: Li, Jingliang, et al.
Published: (2026)
by: Li, Jingliang, et al.
Published: (2026)
Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination
by: Zhong, Xinliu, et al.
Published: (2025)
by: Zhong, Xinliu, et al.
Published: (2025)
Fusion-Mamba for Cross-modality Object Detection
by: Dong, Wenhao, et al.
Published: (2024)
by: Dong, Wenhao, et al.
Published: (2024)
O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation
by: Tian, Tongxuan, et al.
Published: (2025)
by: Tian, Tongxuan, et al.
Published: (2025)
Similar Items
-
Part-Aware Open-Vocabulary 3D Affordance Grounding via Prototypical Semantic and Geometric Alignment
by: Gou, Dongqiang, et al.
Published: (2026) -
Intra-view and Inter-view Correlation Guided Multi-view Novel Class Discovery
by: Wan, Xinhang, et al.
Published: (2025) -
Generalized Deep Multi-view Clustering via Causal Learning with Partially Aligned Cross-view Correspondence
by: Yang, Xihong, et al.
Published: (2025) -
Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion Recognition
by: He, Wen-Jue, et al.
Published: (2025) -
Contrastive Continual Multi-view Clustering with Filtered Structural Fusion
by: Wan, Xinhang, et al.
Published: (2023)