VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Hanqing, Liu, Mingyu, Chen, Xiaoyu, MA, Chengwei, Zhong, Yiming, Yin, Wenti, Liu, Yuhao, Cui, Zhiqing, Yuan, Jiahao, Dai, Lu, Ma, Zhiyuan, Xiong, Hui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
von: Wang, Hanqing, et al.
Veröffentlicht: (2025)
von: Wang, Hanqing, et al.
Veröffentlicht: (2025)
DAG: Unleash the Potential of Diffusion Model for Open-Vocabulary 3D Affordance Grounding
von: Wang, Hanqing, et al.
Veröffentlicht: (2025)
von: Wang, Hanqing, et al.
Veröffentlicht: (2025)
VAGNet: Grounding 3D Affordance from Human-Object Interactions in Videos
von: Mao, Aihua, et al.
Veröffentlicht: (2026)
von: Mao, Aihua, et al.
Veröffentlicht: (2026)
SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model
von: Yu, Chunlin, et al.
Veröffentlicht: (2024)
von: Yu, Chunlin, et al.
Veröffentlicht: (2024)
WorldAfford: Affordance Grounding based on Natural Language Instructions
von: Chen, Changmao, et al.
Veröffentlicht: (2024)
von: Chen, Changmao, et al.
Veröffentlicht: (2024)
O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation
von: Tian, Tongxuan, et al.
Veröffentlicht: (2025)
von: Tian, Tongxuan, et al.
Veröffentlicht: (2025)
PreAfford: Universal Affordance-Based Pre-Grasping for Diverse Objects and Environments
von: Ding, Kairui, et al.
Veröffentlicht: (2024)
von: Ding, Kairui, et al.
Veröffentlicht: (2024)
AffordTissue: Dense Affordance Prediction for Tool-Action Specific Tissue Interaction
von: Maksutova, Aiza, et al.
Veröffentlicht: (2026)
von: Maksutova, Aiza, et al.
Veröffentlicht: (2026)
Afford-X: Generalizable and Slim Affordance Reasoning for Task-oriented Manipulation
von: Zhu, Xiaomeng, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaomeng, et al.
Veröffentlicht: (2025)
Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram Generation
von: Cui, Zhiqing, et al.
Veröffentlicht: (2025)
von: Cui, Zhiqing, et al.
Veröffentlicht: (2025)
AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesis
von: Wu, Xiaofei, et al.
Veröffentlicht: (2026)
von: Wu, Xiaofei, et al.
Veröffentlicht: (2026)
AffordDP: Generalizable Diffusion Policy with Transferable Affordance
von: Wu, Shijie, et al.
Veröffentlicht: (2024)
von: Wu, Shijie, et al.
Veröffentlicht: (2024)
Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
von: Zhu, He, et al.
Veröffentlicht: (2025)
von: Zhu, He, et al.
Veröffentlicht: (2025)
AffordSim: A Scalable Data Generator and Benchmark for Affordance-Aware Robotic Manipulation
von: Li, Mingyang, et al.
Veröffentlicht: (2026)
von: Li, Mingyang, et al.
Veröffentlicht: (2026)
SeqAffordSplat: Scene-level Sequential Affordance Reasoning on 3D Gaussian Splatting
von: Li, Di, et al.
Veröffentlicht: (2025)
von: Li, Di, et al.
Veröffentlicht: (2025)
AffordDexGrasp: Open-set Language-guided Dexterous Grasp with Generalizable-Instructive Affordance
von: Wei, Yi-Lin, et al.
Veröffentlicht: (2025)
von: Wei, Yi-Lin, et al.
Veröffentlicht: (2025)
VoxAfford: Multi-Scale Voxel-Token Fusion for Open-Vocabulary 3D Affordance Detection
von: Sun, Haowen, et al.
Veröffentlicht: (2026)
von: Sun, Haowen, et al.
Veröffentlicht: (2026)
Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding
von: Yang, Zaiquan, et al.
Veröffentlicht: (2025)
von: Yang, Zaiquan, et al.
Veröffentlicht: (2025)
RoboAfford++: A Generative AI-Enhanced Dataset for Multimodal Affordance Learning in Robotic Manipulation and Navigation
von: Hao, Xiaoshuai, et al.
Veröffentlicht: (2025)
von: Hao, Xiaoshuai, et al.
Veröffentlicht: (2025)
EqvAfford: SE(3) Equivariance for Point-Level Affordance Learning
von: Chen, Yue, et al.
Veröffentlicht: (2024)
von: Chen, Yue, et al.
Veröffentlicht: (2024)
Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance
von: Wang, Runze, et al.
Veröffentlicht: (2026)
von: Wang, Runze, et al.
Veröffentlicht: (2026)
One-Shot Affordance Grounding of Deformable Objects in Egocentric Organizing Scenes
von: Jia, Wanjun, et al.
Veröffentlicht: (2025)
von: Jia, Wanjun, et al.
Veröffentlicht: (2025)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers
von: Vu, Nghia, et al.
Veröffentlicht: (2026)
von: Vu, Nghia, et al.
Veröffentlicht: (2026)
RHINO: Reconstructing Human Interactions with Novel Objects from Monocular Videos
von: Xue, Lixin, et al.
Veröffentlicht: (2026)
von: Xue, Lixin, et al.
Veröffentlicht: (2026)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Afford Correspondence
von: Zhang, Jiawei, et al.
Veröffentlicht: (2026)
von: Zhang, Jiawei, et al.
Veröffentlicht: (2026)
AffordanceLLM: Grounding Affordance from Vision Language Models
von: Qian, Shengyi, et al.
Veröffentlicht: (2024)
von: Qian, Shengyi, et al.
Veröffentlicht: (2024)
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
von: Zhou, Donghao, et al.
Veröffentlicht: (2026)
von: Zhou, Donghao, et al.
Veröffentlicht: (2026)
AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter
von: Tang, Yingbo, et al.
Veröffentlicht: (2025)
von: Tang, Yingbo, et al.
Veröffentlicht: (2025)
EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion
von: Wei, Jiangchuan, et al.
Veröffentlicht: (2025)
von: Wei, Jiangchuan, et al.
Veröffentlicht: (2025)
Grounding 3D Scene Affordance From Egocentric Interactions
von: Liu, Cuiyu, et al.
Veröffentlicht: (2024)
von: Liu, Cuiyu, et al.
Veröffentlicht: (2024)
HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance Guidance
von: Li, Lei, et al.
Veröffentlicht: (2025)
von: Li, Lei, et al.
Veröffentlicht: (2025)
Object-Shot Enhanced Grounding Network for Egocentric Video
von: Feng, Yisen, et al.
Veröffentlicht: (2025)
von: Feng, Yisen, et al.
Veröffentlicht: (2025)
Kardia-R1: Unleashing LLMs to Reason toward Understanding and Empathy for Emotional Support via Rubric-as-Judge Reinforcement Learning
von: Yuan, Jiahao, et al.
Veröffentlicht: (2025)
von: Yuan, Jiahao, et al.
Veröffentlicht: (2025)
AffordanceSAM: Segment Anything Once More in Affordance Grounding
von: Jiang, Dengyang, et al.
Veröffentlicht: (2025)
von: Jiang, Dengyang, et al.
Veröffentlicht: (2025)
FSAG: Enhancing Human-to-Dexterous-Hand Finger-Specific Affordance Grounding via Diffusion Models
von: Han, Yifan, et al.
Veröffentlicht: (2026)
von: Han, Yifan, et al.
Veröffentlicht: (2026)
VC-Agent: An Interactive Agent for Customized Video Dataset Collection
von: Zhang, Yidan, et al.
Veröffentlicht: (2025)
von: Zhang, Yidan, et al.
Veröffentlicht: (2025)
PEOD: A Pixel-Aligned Event-RGB Benchmark for Object Detection under Challenging Conditions
von: Cui, Luoping, et al.
Veröffentlicht: (2025)
von: Cui, Luoping, et al.
Veröffentlicht: (2025)
Affordance-First Decomposition for Continual Learning in Video-Language Understanding
von: Xu, Mengzhu, et al.
Veröffentlicht: (2025)
von: Xu, Mengzhu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
von: Wang, Hanqing, et al.
Veröffentlicht: (2025) -
DAG: Unleash the Potential of Diffusion Model for Open-Vocabulary 3D Affordance Grounding
von: Wang, Hanqing, et al.
Veröffentlicht: (2025) -
VAGNet: Grounding 3D Affordance from Human-Object Interactions in Videos
von: Mao, Aihua, et al.
Veröffentlicht: (2026) -
SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model
von: Yu, Chunlin, et al.
Veröffentlicht: (2024) -
WorldAfford: Affordance Grounding based on Natural Language Instructions
von: Chen, Changmao, et al.
Veröffentlicht: (2024)