A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zixin, Chen, Kanghao, Wang, Hanqing, Zhang, Hongfei, Chen, Harold Haodong, Liao, Chenfei, Guo, Litao, Chen, Ying-Cong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Panoramic Affordance Prediction
by: Zhang, Zixin, et al.
Published: (2026)
by: Zhang, Zixin, et al.
Published: (2026)
Show, Don't Tell: Morphing Latent Reasoning into Image Generation
by: Chen, Harold Haodong, et al.
Published: (2026)
by: Chen, Harold Haodong, et al.
Published: (2026)
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
by: Zhang, Zixin, et al.
Published: (2025)
by: Zhang, Zixin, et al.
Published: (2025)
RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data
by: Chen, Harold Haodong, et al.
Published: (2026)
by: Chen, Harold Haodong, et al.
Published: (2026)
DVD: Deterministic Video Depth Estimation with Generative Priors
by: Zhang, Hongfei, et al.
Published: (2026)
by: Zhang, Hongfei, et al.
Published: (2026)
TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
by: Chen, Harold Haodong, et al.
Published: (2025)
by: Chen, Harold Haodong, et al.
Published: (2025)
DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation
by: Zhang, Hongfei, et al.
Published: (2025)
by: Zhang, Hongfei, et al.
Published: (2025)
Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
by: Wang, Hanqing, et al.
Published: (2025)
by: Wang, Hanqing, et al.
Published: (2025)
RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation
by: Kuang, Yuxuan, et al.
Published: (2024)
by: Kuang, Yuxuan, et al.
Published: (2024)
Affordance Agent Harness: Verification-Gated Skill Orchestration
by: Huang, Haojian, et al.
Published: (2026)
by: Huang, Haojian, et al.
Published: (2026)
Elite-EvGS: Learning Event-based 3D Gaussian Splatting by Distilling Event-to-Video Priors
by: Zhang, Zixin, et al.
Published: (2024)
by: Zhang, Zixin, et al.
Published: (2024)
AffordanceLLM: Grounding Affordance from Vision Language Models
by: Qian, Shengyi, et al.
Published: (2024)
by: Qian, Shengyi, et al.
Published: (2024)
DriveVA: Video Action Models are Zero-Shot Drivers
by: Liu, Mengmeng, et al.
Published: (2026)
by: Liu, Mengmeng, et al.
Published: (2026)
One-Shot Affordance Grounding of Deformable Objects in Egocentric Organizing Scenes
by: Jia, Wanjun, et al.
Published: (2025)
by: Jia, Wanjun, et al.
Published: (2025)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
GSMem: 3D Gaussian Splatting as Persistent Spatial Memory for Zero-Shot Embodied Exploration and Reasoning
by: Lu, Yiren, et al.
Published: (2026)
by: Lu, Yiren, et al.
Published: (2026)
RANGER: A Monocular Zero-Shot Semantic Navigation Framework through Visual Contextual Adaptation
by: Yu, Ming-Ming, et al.
Published: (2025)
by: Yu, Ming-Ming, et al.
Published: (2025)
Agentic Scene Policies: Unifying Space, Semantics, and Affordances for Robot Action
by: Morin, Sacha, et al.
Published: (2025)
by: Morin, Sacha, et al.
Published: (2025)
AffordanceGrasp-R1:Leveraging Reasoning-Based Affordance Segmentation with Reinforcement Learning for Robotic Grasping
by: Zhou, Dingyi, et al.
Published: (2026)
by: Zhou, Dingyi, et al.
Published: (2026)
PAVLM: Advancing Point Cloud based Affordance Understanding Via Vision-Language Model
by: Liu, Shang-Ching, et al.
Published: (2024)
by: Liu, Shang-Ching, et al.
Published: (2024)
SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models
by: Zantout, Nader, et al.
Published: (2025)
by: Zantout, Nader, et al.
Published: (2025)
ReasonNavi: Human-Inspired Global Map Reasoning for Zero-Shot Embodied Navigation
by: Ao, Yuzhuo, et al.
Published: (2026)
by: Ao, Yuzhuo, et al.
Published: (2026)
3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D Worlds
by: Chu, Hengshuo, et al.
Published: (2025)
by: Chu, Hengshuo, et al.
Published: (2025)
ZISVFM: Zero-Shot Object Instance Segmentation in Indoor Robotic Environments with Vision Foundation Models
by: Zhang, Ying, et al.
Published: (2025)
by: Zhang, Ying, et al.
Published: (2025)
VoxAfford: Multi-Scale Voxel-Token Fusion for Open-Vocabulary 3D Affordance Detection
by: Sun, Haowen, et al.
Published: (2026)
by: Sun, Haowen, et al.
Published: (2026)
PACA: Perspective-Aware Cross-Attention Representation for Zero-Shot Scene Rearrangement
by: Jin, Shutong, et al.
Published: (2024)
by: Jin, Shutong, et al.
Published: (2024)
PanoNav: Mapless Zero-Shot Object Navigation with Panoramic Scene Parsing and Dynamic Memory
by: Jin, Qunchao, et al.
Published: (2025)
by: Jin, Qunchao, et al.
Published: (2025)
VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation
by: Chen, Hanzhi, et al.
Published: (2025)
by: Chen, Hanzhi, et al.
Published: (2025)
TRACER: Texture-Robust Affordance Chain-of-Thought for Deformable-Object Refinement
by: Jia, Wanjun, et al.
Published: (2026)
by: Jia, Wanjun, et al.
Published: (2026)
RoboPCA: Pose-centered Affordance Learning from Human Demonstrations for Robot Manipulation
by: Xiao, Zhanqi, et al.
Published: (2026)
by: Xiao, Zhanqi, et al.
Published: (2026)
Afford-X: Generalizable and Slim Affordance Reasoning for Task-oriented Manipulation
by: Zhu, Xiaomeng, et al.
Published: (2025)
by: Zhu, Xiaomeng, et al.
Published: (2025)
GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping
by: Ma, Teli, et al.
Published: (2024)
by: Ma, Teli, et al.
Published: (2024)
Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
by: Chen, Kehan, et al.
Published: (2024)
by: Chen, Kehan, et al.
Published: (2024)
Dense Monocular Motion Segmentation Using Optical Flow and Pseudo Depth Map: A Zero-Shot Approach
by: Huang, Yuxiang, et al.
Published: (2024)
by: Huang, Yuxiang, et al.
Published: (2024)
Synthesizing the Kill Chain: A Zero-Shot Framework for Target Verification and Tactical Reasoning on the Edge
by: Barkley, Jesse, et al.
Published: (2026)
by: Barkley, Jesse, et al.
Published: (2026)
Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving
by: Zhang, Dapeng, et al.
Published: (2025)
by: Zhang, Dapeng, et al.
Published: (2025)
DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action
by: Fang, Zhen, et al.
Published: (2025)
by: Fang, Zhen, et al.
Published: (2025)
Egocentric Instruction-oriented Affordance Prediction via Large Multimodal Model
by: Ji, Bokai, et al.
Published: (2025)
by: Ji, Bokai, et al.
Published: (2025)
Zero-Shot UAV Navigation in Forests via Relightable 3D Gaussian Splatting
by: Lv, Zinan, et al.
Published: (2026)
by: Lv, Zinan, et al.
Published: (2026)
PreAfford: Universal Affordance-Based Pre-Grasping for Diverse Objects and Environments
by: Ding, Kairui, et al.
Published: (2024)
by: Ding, Kairui, et al.
Published: (2024)
Similar Items
-
Panoramic Affordance Prediction
by: Zhang, Zixin, et al.
Published: (2026) -
Show, Don't Tell: Morphing Latent Reasoning into Image Generation
by: Chen, Harold Haodong, et al.
Published: (2026) -
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
by: Zhang, Zixin, et al.
Published: (2025) -
RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data
by: Chen, Harold Haodong, et al.
Published: (2026) -
DVD: Deterministic Video Depth Estimation with Generative Priors
by: Zhang, Hongfei, et al.
Published: (2026)