WorldAfford: Affordance Grounding based on Natural Language Instructions
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Changmao, Cong, Yuren, Kan, Zhen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model
by: Yu, Chunlin, et al.
Published: (2024)
by: Yu, Chunlin, et al.
Published: (2024)
Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
by: Zhu, He, et al.
Published: (2025)
by: Zhu, He, et al.
Published: (2025)
AffordTissue: Dense Affordance Prediction for Tool-Action Specific Tissue Interaction
by: Maksutova, Aiza, et al.
Published: (2026)
by: Maksutova, Aiza, et al.
Published: (2026)
Selective Contrastive Learning for Weakly Supervised Affordance Grounding
by: Moon, WonJun, et al.
Published: (2025)
by: Moon, WonJun, et al.
Published: (2025)
Grounding 3D Scene Affordance From Egocentric Interactions
by: Liu, Cuiyu, et al.
Published: (2024)
by: Liu, Cuiyu, et al.
Published: (2024)
CLIPDrag: Combining Text-based and Drag-based Instructions for Image Editing
by: Jiang, Ziqi, et al.
Published: (2024)
by: Jiang, Ziqi, et al.
Published: (2024)
GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding
by: Shao, Yawen, et al.
Published: (2024)
by: Shao, Yawen, et al.
Published: (2024)
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
by: Pei, Siqi, et al.
Published: (2026)
by: Pei, Siqi, et al.
Published: (2026)
A Large-scale Dataset for Robust Complex Anime Scene Text Detection
by: Dong, Ziyi, et al.
Published: (2025)
by: Dong, Ziyi, et al.
Published: (2025)
BEACON: Language-Conditioned Navigation Affordance Prediction under Occlusion
by: Gao, Xinyu, et al.
Published: (2026)
by: Gao, Xinyu, et al.
Published: (2026)
Vega: Learning to Drive with Natural Language Instructions
by: Zuo, Sicheng, et al.
Published: (2026)
by: Zuo, Sicheng, et al.
Published: (2026)
OpenGround: Active Cognition-based Reasoning for Open-World 3D Visual Grounding
by: Huang, Wenyuan, et al.
Published: (2025)
by: Huang, Wenyuan, et al.
Published: (2025)
VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model
by: Wang, Hanqing, et al.
Published: (2026)
by: Wang, Hanqing, et al.
Published: (2026)
Can Segmentation Models Understand the World? Towards Proactive Affordance Reasoning via Visual Chain-of-Thought
by: Guo, Yuchen, et al.
Published: (2026)
by: Guo, Yuchen, et al.
Published: (2026)
Segment Any Object Model (SAOM): Real-to-Simulation Fine-Tuning Strategy for Multi-Class Multi-Instance Segmentation
by: Khan, Mariia, et al.
Published: (2024)
by: Khan, Mariia, et al.
Published: (2024)
UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction-as-Reasoning
by: Chen, Liangyu, et al.
Published: (2025)
by: Chen, Liangyu, et al.
Published: (2025)
PhyDrawGen: Physically Grounded Diagram Generation from Natural Language
by: Haque, Nafiul, et al.
Published: (2026)
by: Haque, Nafiul, et al.
Published: (2026)
doScenes: An Autonomous Driving Dataset with Natural Language Instruction for Human Interaction and Vision-Language Navigation
by: Roy, Parthib, et al.
Published: (2024)
by: Roy, Parthib, et al.
Published: (2024)
Script-to-Slide Grounding: Grounding Script Sentences to Slide Objects for Automatic Instructional Video Generation
by: Suzuki, Rena, et al.
Published: (2026)
by: Suzuki, Rena, et al.
Published: (2026)
Self-Explainable Affordance Learning with Embodied Caption
by: Zhang, Zhipeng, et al.
Published: (2024)
by: Zhang, Zhipeng, et al.
Published: (2024)
SYNTHIA: Novel Concept Design with Affordance Composition
by: Ha, Hyeonjeong, et al.
Published: (2025)
by: Ha, Hyeonjeong, et al.
Published: (2025)
MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation
by: Zhang, Pingrui, et al.
Published: (2025)
by: Zhang, Pingrui, et al.
Published: (2025)
Text-driven Affordance Learning from Egocentric Vision
by: Yoshida, Tomoya, et al.
Published: (2024)
by: Yoshida, Tomoya, et al.
Published: (2024)
Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding
by: Yao, Yuan, et al.
Published: (2026)
by: Yao, Yuan, et al.
Published: (2026)
EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models
by: Villa, Andrés, et al.
Published: (2025)
by: Villa, Andrés, et al.
Published: (2025)
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
by: Yuan, Wentao, et al.
Published: (2024)
by: Yuan, Wentao, et al.
Published: (2024)
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy
by: Yang, Te, et al.
Published: (2024)
by: Yang, Te, et al.
Published: (2024)
World-To-Image: Grounding Text-to-Image Generation with Agent-Driven World Knowledge
by: Son, Moo Hyun, et al.
Published: (2025)
by: Son, Moo Hyun, et al.
Published: (2025)
GroundCount: Grounding Vision-Language Models with Object Detection for Mitigating Counting Hallucinations
by: Chen, Boyuan, et al.
Published: (2026)
by: Chen, Boyuan, et al.
Published: (2026)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
by: Qin, Bosheng, et al.
Published: (2023)
by: Qin, Bosheng, et al.
Published: (2023)
VIMI: Grounding Video Generation through Multi-modal Instruction
by: Fang, Yuwei, et al.
Published: (2024)
by: Fang, Yuwei, et al.
Published: (2024)
OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios
by: Gao, Hong, et al.
Published: (2025)
by: Gao, Hong, et al.
Published: (2025)
Capturing Fine-Grained Alignments Improves 3D Affordance Detection
by: Tokumitsu, Junsei, et al.
Published: (2025)
by: Tokumitsu, Junsei, et al.
Published: (2025)
Informative Text-Image Alignment for Visual Affordance Learning with Foundation Models
by: Zhang, Qian, et al.
Published: (2025)
by: Zhang, Qian, et al.
Published: (2025)
MOAT: Evaluating LMMs for Capability Integration and Instruction Grounding
by: Ye, Zhoutong, et al.
Published: (2025)
by: Ye, Zhoutong, et al.
Published: (2025)
O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation
by: Tian, Tongxuan, et al.
Published: (2025)
by: Tian, Tongxuan, et al.
Published: (2025)
MTA-RL: Robust Urban Driving via Multi-modal Transformer-based 3D Affordances and Reinforcement Learning
by: Chen, Guangli, et al.
Published: (2026)
by: Chen, Guangli, et al.
Published: (2026)
AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers
by: Vu, Nghia, et al.
Published: (2026)
by: Vu, Nghia, et al.
Published: (2026)
Large Language Model with Region-guided Referring and Grounding for CT Report Generation
by: Chen, Zhixuan, et al.
Published: (2024)
by: Chen, Zhixuan, et al.
Published: (2024)
Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles
by: Liao, Haicheng, et al.
Published: (2025)
by: Liao, Haicheng, et al.
Published: (2025)
Similar Items
-
SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model
by: Yu, Chunlin, et al.
Published: (2024) -
Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
by: Zhu, He, et al.
Published: (2025) -
AffordTissue: Dense Affordance Prediction for Tool-Action Specific Tissue Interaction
by: Maksutova, Aiza, et al.
Published: (2026) -
Selective Contrastive Learning for Weakly Supervised Affordance Grounding
by: Moon, WonJun, et al.
Published: (2025) -
Grounding 3D Scene Affordance From Egocentric Interactions
by: Liu, Cuiyu, et al.
Published: (2024)