Visual-Geometric Collaborative Guidance for Affordance Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Hongchen, Zhai, Wei, Wang, Jiao, Cao, Yang, Zha, Zheng-Jun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leverage Task Context for Object Affordance Ranking
by: Huang, Haojie, et al.
Published: (2024)
by: Huang, Haojie, et al.
Published: (2024)
GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding
by: Shao, Yawen, et al.
Published: (2024)
by: Shao, Yawen, et al.
Published: (2024)
Grounding 3D Scene Affordance From Egocentric Interactions
by: Liu, Cuiyu, et al.
Published: (2024)
by: Liu, Cuiyu, et al.
Published: (2024)
LEMON: Learning 3D Human-Object Interaction Relation from 2D Images
by: Yang, Yuhang, et al.
Published: (2023)
by: Yang, Yuhang, et al.
Published: (2023)
VMAD: Visual-enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection
by: Deng, Huilin, et al.
Published: (2024)
by: Deng, Huilin, et al.
Published: (2024)
Intention-driven Ego-to-Exo Video Generation
by: Luo, Hongchen, et al.
Published: (2024)
by: Luo, Hongchen, et al.
Published: (2024)
Event-based Visual Deformation Measurement
by: Wu, Yuliang, et al.
Published: (2026)
by: Wu, Yuliang, et al.
Published: (2026)
Bidirectional Progressive Transformer for Interaction Intention Anticipation
by: Zhang, Zichen, et al.
Published: (2024)
by: Zhang, Zichen, et al.
Published: (2024)
PEAR: Phrase-Based Hand-Object Interaction Anticipation
by: Zhang, Zichen, et al.
Published: (2024)
by: Zhang, Zichen, et al.
Published: (2024)
HERO: Human Reaction Generation from Videos
by: Yu, Chengjun, et al.
Published: (2025)
by: Yu, Chengjun, et al.
Published: (2025)
End-to-End Spatial-Temporal Transformer for Real-time 4D HOI Reconstruction
by: Zhang, Haoyu, et al.
Published: (2026)
by: Zhang, Haoyu, et al.
Published: (2026)
TOUCH: Text-guided Controllable Generation of Free-Form Hand-Object Interactions
by: Han, Guangyi, et al.
Published: (2025)
by: Han, Guangyi, et al.
Published: (2025)
Event Stream Filtering via Probability Flux Estimation
by: Chen, Jinze, et al.
Published: (2025)
by: Chen, Jinze, et al.
Published: (2025)
EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric Views
by: Yang, Yuhang, et al.
Published: (2024)
by: Yang, Yuhang, et al.
Published: (2024)
Learning Visual Affordance from Audio
by: Lu, Lidong, et al.
Published: (2025)
by: Lu, Lidong, et al.
Published: (2025)
Visual Context Window Extension: A New Perspective for Long Video Understanding
by: Wei, Hongchen, et al.
Published: (2024)
by: Wei, Hongchen, et al.
Published: (2024)
OmniVaT: Single Domain Generalization for Multimodal Visual-Tactile Learning
by: Qiu, Liuxiang, et al.
Published: (2026)
by: Qiu, Liuxiang, et al.
Published: (2026)
EMoTive: Event-guided Trajectory Modeling for 3D Motion Estimation
by: Wan, Zengyu, et al.
Published: (2025)
by: Wan, Zengyu, et al.
Published: (2025)
GRACE: Estimating Geometry-level 3D Human-Scene Contact from 2D Images
by: Wang, Chengfeng, et al.
Published: (2025)
by: Wang, Chengfeng, et al.
Published: (2025)
Towards Better De-raining Generalization via Rainy Characteristics Memorization and Replay
by: Wang, Kunyu, et al.
Published: (2025)
by: Wang, Kunyu, et al.
Published: (2025)
MATE: Motion-Augmented Temporal Consistency for Event-based Point Tracking
by: Han, Han, et al.
Published: (2024)
by: Han, Han, et al.
Published: (2024)
SafeCFG: Controlling Harmful Features with Dynamic Safe Guidance for Safe Generation
by: Pan, Jiadong, et al.
Published: (2024)
by: Pan, Jiadong, et al.
Published: (2024)
MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling
by: Yang, Jian, et al.
Published: (2024)
by: Yang, Jian, et al.
Published: (2024)
ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs
by: Gao, Yiling, et al.
Published: (2026)
by: Gao, Yiling, et al.
Published: (2026)
MatE: Material Extraction from Single-Image via Geometric Prior
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
ViewPoint: Panoramic Video Generation with Pretrained Diffusion Models
by: Fang, Zixun, et al.
Published: (2025)
by: Fang, Zixun, et al.
Published: (2025)
Unbiased Gradient Estimation for Event Binning via Functional Backpropagation
by: Chen, Jinze, et al.
Published: (2026)
by: Chen, Jinze, et al.
Published: (2026)
Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning
by: Deng, Huilin, et al.
Published: (2025)
by: Deng, Huilin, et al.
Published: (2025)
PAID: Pairwise Angular-Invariant Decomposition for Continual Test-Time Adaptation
by: Wang, Kunyu, et al.
Published: (2025)
by: Wang, Kunyu, et al.
Published: (2025)
Efficient Test-time Adaptive Object Detection via Sensitivity-Guided Pruning
by: Wang, Kunyu, et al.
Published: (2025)
by: Wang, Kunyu, et al.
Published: (2025)
EF-3DGS: Event-Aided Free-Trajectory 3D Gaussian Splatting
by: Liao, Bohao, et al.
Published: (2024)
by: Liao, Bohao, et al.
Published: (2024)
Gloria: Consistent Character Video Generation via Content Anchors
by: Yang, Yuhang, et al.
Published: (2026)
by: Yang, Yuhang, et al.
Published: (2026)
Closed-Loop Transfer for Weakly-supervised Affordance Grounding
by: Tang, Jiajin, et al.
Published: (2025)
by: Tang, Jiajin, et al.
Published: (2025)
AffordanceSAM: Segment Anything Once More in Affordance Grounding
by: Jiang, Dengyang, et al.
Published: (2025)
by: Jiang, Dengyang, et al.
Published: (2025)
Training-Free Reasoning and Reflection in MLLMs
by: Wei, Hongchen, et al.
Published: (2025)
by: Wei, Hongchen, et al.
Published: (2025)
FC3DNet: A Fully Connected Encoder-Decoder for Efficient Demoir'eing
by: Du, Zhibo, et al.
Published: (2024)
by: Du, Zhibo, et al.
Published: (2024)
RAIN: Real-time Animation of Infinite Video Stream
by: Shu, Zhilei, et al.
Published: (2024)
by: Shu, Zhilei, et al.
Published: (2024)
Implicit and Explicit Language Guidance for Diffusion-based Visual Perception
by: Wang, Hefeng, et al.
Published: (2024)
by: Wang, Hefeng, et al.
Published: (2024)
Learning 2D Invariant Affordance Knowledge for 3D Affordance Grounding
by: Gao, Xianqiang, et al.
Published: (2024)
by: Gao, Xianqiang, et al.
Published: (2024)
RCGNet: RGB-based Category-Level 6D Object Pose Estimation with Geometric Guidance
by: Yu, Sheng, et al.
Published: (2025)
by: Yu, Sheng, et al.
Published: (2025)
Similar Items
-
Leverage Task Context for Object Affordance Ranking
by: Huang, Haojie, et al.
Published: (2024) -
GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding
by: Shao, Yawen, et al.
Published: (2024) -
Grounding 3D Scene Affordance From Egocentric Interactions
by: Liu, Cuiyu, et al.
Published: (2024) -
LEMON: Learning 3D Human-Object Interaction Relation from 2D Images
by: Yang, Yuhang, et al.
Published: (2023) -
VMAD: Visual-enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection
by: Deng, Huilin, et al.
Published: (2024)