Mitigating Cross-Modal Distraction and Ensuring Geometric Feasibility via Affordance-Guided and Self-Consistent MLLMs for Task Planning in Instruction-Following Manipulation
Fuente:
arXiv
Guardado en:
| Autores principales: | Shen, Yu-Hong, Wu, Chuan-Yu, Yang, Yi-Ru, Tai, Yen-Ling, Chen, Yi-Ting |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
GRITS: A Spillage-Aware Guided Diffusion Policy for Robot Food Scooping Tasks
por: Tai, Yen-Ling, et al.
Publicado: (2025)
por: Tai, Yen-Ling, et al.
Publicado: (2025)
Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge
por: Chen, Yan-Lun, et al.
Publicado: (2025)
por: Chen, Yan-Lun, et al.
Publicado: (2025)
Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile Manipulation
por: Lin, Tzu-Jung, et al.
Publicado: (2025)
por: Lin, Tzu-Jung, et al.
Publicado: (2025)
Learning Instruction-Guided Manipulation Affordance via Large Models for Embodied Robotic Tasks
por: Li, Dayou, et al.
Publicado: (2024)
por: Li, Dayou, et al.
Publicado: (2024)
Empowering Reliable Visual-Centric Instruction Following in MLLMs
por: He, Weilei, et al.
Publicado: (2026)
por: He, Weilei, et al.
Publicado: (2026)
Self-Guided Plan Extraction for Instruction-Following Tasks with Goal-Conditional Reinforcement Learning
por: Volovikova, Zoya, et al.
Publicado: (2026)
por: Volovikova, Zoya, et al.
Publicado: (2026)
Affordance Benchmark for MLLMs
por: Wang, Junying, et al.
Publicado: (2025)
por: Wang, Junying, et al.
Publicado: (2025)
GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency
por: Lu, Dongyue, et al.
Publicado: (2024)
por: Lu, Dongyue, et al.
Publicado: (2024)
Articulated Object Manipulation with Coarse-to-fine Affordance for Mitigating the Effect of Point Cloud Noise
por: Ling, Suhan, et al.
Publicado: (2024)
por: Ling, Suhan, et al.
Publicado: (2024)
AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesis
por: Wu, Xiaofei, et al.
Publicado: (2026)
por: Wu, Xiaofei, et al.
Publicado: (2026)
O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation
por: Tian, Tongxuan, et al.
Publicado: (2025)
por: Tian, Tongxuan, et al.
Publicado: (2025)
Affordance Design for Mitigating Digital Vulnerability in Smart Tourism
por: Yuanyuan Shi, et al.
Publicado: (2025)
por: Yuanyuan Shi, et al.
Publicado: (2025)
PLATO: Planning with LLMs and Affordances for Tool Manipulation
por: Car, Arvind, et al.
Publicado: (2024)
por: Car, Arvind, et al.
Publicado: (2024)
Task-Aware Bimanual Affordance Prediction via VLM-Guided Semantic-Geometric Reasoning
por: Hahne, Fabian, et al.
Publicado: (2026)
por: Hahne, Fabian, et al.
Publicado: (2026)
ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints
por: Chen, Pei-An, et al.
Publicado: (2026)
por: Chen, Pei-An, et al.
Publicado: (2026)
3D Gaussian Inpainting with Depth-Guided Cross-View Consistency
por: Huang, Sheng-Yu, et al.
Publicado: (2025)
por: Huang, Sheng-Yu, et al.
Publicado: (2025)
Feasibility-Guided Planning over Multi-Specialized Locomotion Policies
por: Luo, Ying-Sheng, et al.
Publicado: (2026)
por: Luo, Ying-Sheng, et al.
Publicado: (2026)
GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance Guidance
por: Yuan, Shuaihang, et al.
Publicado: (2024)
por: Yuan, Shuaihang, et al.
Publicado: (2024)
Distribution-Consistency-Guided Multi-modal Hashing
por: Liu, Jin-Yu, et al.
Publicado: (2024)
por: Liu, Jin-Yu, et al.
Publicado: (2024)
Ensuring Consistency for In-Image Translation
por: Fu, Chengpeng, et al.
Publicado: (2024)
por: Fu, Chengpeng, et al.
Publicado: (2024)
Beyond Unimodal Shortcuts: MLLMs as Cross-Modal Reasoners for Grounded Named Entity Recognition
por: Ma, Jinlong, et al.
Publicado: (2026)
por: Ma, Jinlong, et al.
Publicado: (2026)
Task-Aware 3D Affordance Segmentation via 2D Guidance and Geometric Refinement
por: He, Lian, et al.
Publicado: (2025)
por: He, Lian, et al.
Publicado: (2025)
HAMMER: Harnessing MLLM via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding
por: Yao, Lei, et al.
Publicado: (2026)
por: Yao, Lei, et al.
Publicado: (2026)
GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation
por: Gao, Ning, et al.
Publicado: (2025)
por: Gao, Ning, et al.
Publicado: (2025)
SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs
por: Li, Haoxuan, et al.
Publicado: (2025)
por: Li, Haoxuan, et al.
Publicado: (2025)
Incorporating Task Progress Knowledge for Subgoal Generation in Robotic Manipulation through Image Edits
por: Kang, Xuhui, et al.
Publicado: (2024)
por: Kang, Xuhui, et al.
Publicado: (2024)
DistractMIA: Black-Box Membership Inference on Vision-Language Models via Semantic Distraction
por: Tang, Hongyi, et al.
Publicado: (2026)
por: Tang, Hongyi, et al.
Publicado: (2026)
A Hidden Stumbling Block in Generalized Category Discovery: Distracted Attention
por: Xu, Qiyu, et al.
Publicado: (2025)
por: Xu, Qiyu, et al.
Publicado: (2025)
AFFORD2ACT: Affordance-Guided Automatic Keypoint Selection for Generalizable and Lightweight Robotic Manipulation
por: Singh, Anukriti, et al.
Publicado: (2025)
por: Singh, Anukriti, et al.
Publicado: (2025)
Octavius: Mitigating Task Interference in MLLMs via LoRA-MoE
por: Chen, Zeren, et al.
Publicado: (2023)
por: Chen, Zeren, et al.
Publicado: (2023)
NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model
por: Lin, Yen-Ting, et al.
Publicado: (2024)
por: Lin, Yen-Ting, et al.
Publicado: (2024)
SG-XDEAT: Sparsity-Guided Cross-Dimensional and Cross-Encoding Attention with Target-Aware Conditioning in Tabular Learning
por: Cheng, Chih-Chuan, et al.
Publicado: (2025)
por: Cheng, Chih-Chuan, et al.
Publicado: (2025)
Afford-X: Generalizable and Slim Affordance Reasoning for Task-oriented Manipulation
por: Zhu, Xiaomeng, et al.
Publicado: (2025)
por: Zhu, Xiaomeng, et al.
Publicado: (2025)
Mitigating Spurious Correlations in Weakly Supervised Semantic Segmentation via Cross-architecture Consistency Regularization
por: Zhang, Zheyuan, et al.
Publicado: (2025)
por: Zhang, Zheyuan, et al.
Publicado: (2025)
Residual Cross-Modal Fusion Networks for Audio-Visual Navigation
por: Wang, Yi, et al.
Publicado: (2026)
por: Wang, Yi, et al.
Publicado: (2026)
Tip to Overcome Oversized Finger Traps in Wrist Arthroscopy Distraction
por: I‐Ning Lo, et al.
Publicado: (2025)
por: I‐Ning Lo, et al.
Publicado: (2025)
Multi-Modal Face Anti-Spoofing via Cross-Modal Feature Transitions
por: Chong, Jun-Xiong, et al.
Publicado: (2025)
por: Chong, Jun-Xiong, et al.
Publicado: (2025)
DORA: Object Affordance-Guided Reinforcement Learning for Dexterous Robotic Manipulation
por: Zhang, Lei, et al.
Publicado: (2025)
por: Zhang, Lei, et al.
Publicado: (2025)
LLMs can be easily Confused by Instructional Distractions
por: Hwang, Yerin, et al.
Publicado: (2025)
por: Hwang, Yerin, et al.
Publicado: (2025)
RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
por: Nasiriany, Soroush, et al.
Publicado: (2024)
por: Nasiriany, Soroush, et al.
Publicado: (2024)
Ejemplares similares
-
GRITS: A Spillage-Aware Guided Diffusion Policy for Robot Food Scooping Tasks
por: Tai, Yen-Ling, et al.
Publicado: (2025) -
Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge
por: Chen, Yan-Lun, et al.
Publicado: (2025) -
Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile Manipulation
por: Lin, Tzu-Jung, et al.
Publicado: (2025) -
Learning Instruction-Guided Manipulation Affordance via Large Models for Embodied Robotic Tasks
por: Li, Dayou, et al.
Publicado: (2024) -
Empowering Reliable Visual-Centric Instruction Following in MLLMs
por: He, Weilei, et al.
Publicado: (2026)