H2OFlow: Grounding Human-Object Affordances with 3D Generative Models and Dense Diffused Flows
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Harry, Carlone, Luca |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CHAMP: Conformalized 3D Human Multi-Hypothesis Pose Estimators
by: Zhang, Harry, et al.
Published: (2024)
by: Zhang, Harry, et al.
Published: (2024)
CUPS: Improving Human Pose-Shape Estimators with Conformalized Deep Uncertainty
by: Zhang, Harry, et al.
Published: (2024)
by: Zhang, Harry, et al.
Published: (2024)
Multi-Model 3D Registration: Finding Multiple Moving Objects in Cluttered Point Clouds
by: Jin, David, et al.
Published: (2024)
by: Jin, David, et al.
Published: (2024)
VAGNet: Grounding 3D Affordance from Human-Object Interactions in Videos
by: Mao, Aihua, et al.
Published: (2026)
by: Mao, Aihua, et al.
Published: (2026)
VGGT-SLAM 2.0: Real-time Dense Feed-forward Scene Reconstruction
by: Maggio, Dominic, et al.
Published: (2026)
by: Maggio, Dominic, et al.
Published: (2026)
Learning 2D Invariant Affordance Knowledge for 3D Affordance Grounding
by: Gao, Xianqiang, et al.
Published: (2024)
by: Gao, Xianqiang, et al.
Published: (2024)
DAG: Unleash the Potential of Diffusion Model for Open-Vocabulary 3D Affordance Grounding
by: Wang, Hanqing, et al.
Published: (2025)
by: Wang, Hanqing, et al.
Published: (2025)
VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold
by: Maggio, Dominic, et al.
Published: (2025)
by: Maggio, Dominic, et al.
Published: (2025)
CRISP: Object Pose and Shape Estimation with Test-Time Adaptation
by: Shi, Jingnan, et al.
Published: (2024)
by: Shi, Jingnan, et al.
Published: (2024)
Affostruction: 3D Affordance Grounding with Generative Reconstruction
by: Park, Chunghyun, et al.
Published: (2026)
by: Park, Chunghyun, et al.
Published: (2026)
VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model
by: Wang, Hanqing, et al.
Published: (2026)
by: Wang, Hanqing, et al.
Published: (2026)
DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-trained Video Diffusion Models
by: Kim, Hyeonwoo, et al.
Published: (2025)
by: Kim, Hyeonwoo, et al.
Published: (2025)
Beyond the Contact: Discovering Comprehensive Affordance for 3D Objects from Pre-trained 2D Diffusion Models
by: Kim, Hyeonwoo, et al.
Published: (2024)
by: Kim, Hyeonwoo, et al.
Published: (2024)
Uncertainty Quantification for Visual Object Pose Estimation
by: Shaikewitz, Lorenzo, et al.
Published: (2025)
by: Shaikewitz, Lorenzo, et al.
Published: (2025)
O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation
by: Tian, Tongxuan, et al.
Published: (2025)
by: Tian, Tongxuan, et al.
Published: (2025)
CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects
by: Li, Jingliang, et al.
Published: (2026)
by: Li, Jingliang, et al.
Published: (2026)
Mixed Diffusion for 3D Indoor Scene Synthesis
by: Hu, Siyi, et al.
Published: (2024)
by: Hu, Siyi, et al.
Published: (2024)
Category-Level Object Shape and Pose Estimation in Less Than a Millisecond
by: Shaikewitz, Lorenzo, et al.
Published: (2025)
by: Shaikewitz, Lorenzo, et al.
Published: (2025)
IAAO: Interactive Affordance Learning for Articulated Objects in 3D Environments
by: Zhang, Can, et al.
Published: (2025)
by: Zhang, Can, et al.
Published: (2025)
Improving 2D-3D Dense Correspondences with Diffusion Models for 6D Object Pose Estimation
by: Hönig, Peter, et al.
Published: (2024)
by: Hönig, Peter, et al.
Published: (2024)
Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
by: Zhu, He, et al.
Published: (2025)
by: Zhu, He, et al.
Published: (2025)
Bayesian Fields: Task-driven Open-Set Semantic Gaussian Splatting
by: Maggio, Dominic, et al.
Published: (2025)
by: Maggio, Dominic, et al.
Published: (2025)
AffordanceSAM: Segment Anything Once More in Affordance Grounding
by: Jiang, Dengyang, et al.
Published: (2025)
by: Jiang, Dengyang, et al.
Published: (2025)
AffordanceLLM: Grounding Affordance from Vision Language Models
by: Qian, Shengyi, et al.
Published: (2024)
by: Qian, Shengyi, et al.
Published: (2024)
GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding
by: Shao, Yawen, et al.
Published: (2024)
by: Shao, Yawen, et al.
Published: (2024)
FlowBot3D: Learning 3D Articulation Flow to Manipulate Articulated Objects
by: Eisner, Ben, et al.
Published: (2022)
by: Eisner, Ben, et al.
Published: (2022)
DiffCLIP: Leveraging Stable Diffusion for Language Grounded 3D Classification
by: Shen, Sitian, et al.
Published: (2023)
by: Shen, Sitian, et al.
Published: (2023)
Pandora: Articulated 3D Scene Graphs from Egocentric Vision
by: Yu, Alan, et al.
Published: (2026)
by: Yu, Alan, et al.
Published: (2026)
Grounding by Remembering: Cross-Scene and In-Scene Memory for 3D Functional Affordances
by: Wang, Qirui, et al.
Published: (2026)
by: Wang, Qirui, et al.
Published: (2026)
SceneTeract: Agentic Functional Affordances and VLM Grounding in 3D Scenes
by: Maillard, Léopold, et al.
Published: (2026)
by: Maillard, Léopold, et al.
Published: (2026)
Affordance-Guided Diffusion Prior for 3D Hand Reconstruction
by: Suzuki, Naru, et al.
Published: (2025)
by: Suzuki, Naru, et al.
Published: (2025)
Object Affordance Recognition and Grounding via Multi-scale Cross-modal Representation Learning
by: Wan, Xinhang, et al.
Published: (2025)
by: Wan, Xinhang, et al.
Published: (2025)
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding
by: Zheng, Henry, et al.
Published: (2025)
by: Zheng, Henry, et al.
Published: (2025)
Affordance-Aware Object Insertion via Mask-Aware Dual Diffusion
by: He, Jixuan, et al.
Published: (2024)
by: He, Jixuan, et al.
Published: (2024)
Grounding 3D Scene Affordance From Egocentric Interactions
by: Liu, Cuiyu, et al.
Published: (2024)
by: Liu, Cuiyu, et al.
Published: (2024)
HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance Guidance
by: Li, Lei, et al.
Published: (2025)
by: Li, Lei, et al.
Published: (2025)
InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance Parsing
by: Zhang, Jinlu, et al.
Published: (2025)
by: Zhang, Jinlu, et al.
Published: (2025)
Populate-A-Scene: Affordance-Aware Human Video Generation
by: Shan, Mengyi, et al.
Published: (2025)
by: Shan, Mengyi, et al.
Published: (2025)
RDM: Recurrent Diffusion Model for Human Motion Generation
by: Mohamed, Mirgahney, et al.
Published: (2024)
by: Mohamed, Mirgahney, et al.
Published: (2024)
MVDiffusion++: A Dense High-resolution Multi-view Diffusion Model for Single or Sparse-view 3D Object Reconstruction
by: Tang, Shitao, et al.
Published: (2024)
by: Tang, Shitao, et al.
Published: (2024)
Similar Items
-
CHAMP: Conformalized 3D Human Multi-Hypothesis Pose Estimators
by: Zhang, Harry, et al.
Published: (2024) -
CUPS: Improving Human Pose-Shape Estimators with Conformalized Deep Uncertainty
by: Zhang, Harry, et al.
Published: (2024) -
Multi-Model 3D Registration: Finding Multiple Moving Objects in Cluttered Point Clouds
by: Jin, David, et al.
Published: (2024) -
VAGNet: Grounding 3D Affordance from Human-Object Interactions in Videos
by: Mao, Aihua, et al.
Published: (2026) -
VGGT-SLAM 2.0: Real-time Dense Feed-forward Scene Reconstruction
by: Maggio, Dominic, et al.
Published: (2026)