Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Bai, Yongjie, Wang, Zhouxia, Liu, Yang, Luo, Kaijun, Wen, Yifan, Dai, Mingtong, Chen, Weixing, Chen, Ziliang, Liu, Lingbo, Li, Guanbin, Lin, Liang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RoVer: Robot Reward Model as Test-Time Verifier for Vision-Language-Action Model
by: Dai, Mingtong, et al.
Published: (2025)
by: Dai, Mingtong, et al.
Published: (2025)
SkiP: When to Skip and When to Refine for Efficient Robot Manipulation
by: Dai, Mingtong, et al.
Published: (2026)
by: Dai, Mingtong, et al.
Published: (2026)
Beyond the Destination: A Novel Benchmark for Exploration-Aware Embodied Question Answering
by: Jiang, Kaixuan, et al.
Published: (2025)
by: Jiang, Kaixuan, et al.
Published: (2025)
GraspView: Active Perception Scoring and Best-View Optimization for Robotic Grasping in Cluttered Environments
by: Wang, Shenglin, et al.
Published: (2025)
by: Wang, Shenglin, et al.
Published: (2025)
Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation
by: Liu, Zixian, et al.
Published: (2025)
by: Liu, Zixian, et al.
Published: (2025)
ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks
by: Wang, Kaijun, et al.
Published: (2025)
by: Wang, Kaijun, et al.
Published: (2025)
Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method
by: Song, Xinshuai, et al.
Published: (2024)
by: Song, Xinshuai, et al.
Published: (2024)
Exploration and Comparison: Development and Implementation of Multiple Ultrasound Imaging Modalities
by: Chen, Xuyang, et al.
Published: (2025)
by: Chen, Xuyang, et al.
Published: (2025)
DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering
by: Luo, Jingzhou, et al.
Published: (2025)
by: Luo, Jingzhou, et al.
Published: (2025)
MEIA: Multimodal Embodied Perception and Interaction in Unknown Environments
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
3DAffordSplat: Efficient Affordance Reasoning with 3D Gaussians
by: Wei, Zeming, et al.
Published: (2025)
by: Wei, Zeming, et al.
Published: (2025)
DDP-WM: Disentangled Dynamics Prediction for Efficient World Models
by: Yin, Shicheng, et al.
Published: (2026)
by: Yin, Shicheng, et al.
Published: (2026)
PhyScene3D: Physically Consistent Interactive 3D Tabletop Scene Generation
by: Chen, Weixing, et al.
Published: (2026)
by: Chen, Weixing, et al.
Published: (2026)
RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models
by: Luo, Jingzhou, et al.
Published: (2026)
by: Luo, Jingzhou, et al.
Published: (2026)
See, Plan, Rewind: Progress-Aware Vision-Language-Action Models for Robust Robotic Manipulation
by: Dai, Tingjun, et al.
Published: (2026)
by: Dai, Tingjun, et al.
Published: (2026)
Cross-Modal Causal Intervention for Medical Report Generation
by: Chen, Weixing, et al.
Published: (2023)
by: Chen, Weixing, et al.
Published: (2023)
PerAct2: Benchmarking and Learning for Robotic Bimanual Manipulation Tasks
by: Grotz, Markus, et al.
Published: (2024)
by: Grotz, Markus, et al.
Published: (2024)
Energy consumption optimization and self-powered environmental monitoring design for low-carbon smart buildings
by: Dai, Yuhan, et al.
Published: (2025)
by: Dai, Yuhan, et al.
Published: (2025)
VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis
by: Gu, Songen, et al.
Published: (2026)
by: Gu, Songen, et al.
Published: (2026)
From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation
by: Yuan, Yifu, et al.
Published: (2025)
by: Yuan, Yifu, et al.
Published: (2025)
Task-Aware Exploration via a Predictive Bisimulation Metric
by: Liang, Dayang, et al.
Published: (2026)
by: Liang, Dayang, et al.
Published: (2026)
PIMbot: A Self-Adaptive Attack Framework for Adversarial Manipulation of Multi-Robot Reinforcement Learning
by: Li, Zexin, et al.
Published: (2026)
by: Li, Zexin, et al.
Published: (2026)
VisualActBench: Can VLMs See and Act like a Human?
by: Zhang, Daoan, et al.
Published: (2025)
by: Zhang, Daoan, et al.
Published: (2025)
Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation
by: Wang, Guokang, et al.
Published: (2024)
by: Wang, Guokang, et al.
Published: (2024)
VTON 360: High-Fidelity Virtual Try-On from Any Viewing Direction
by: He, Zijian, et al.
Published: (2025)
by: He, Zijian, et al.
Published: (2025)
Nd3+ Doping-induced Leakage Currents Suppression in High-temperature 0.7BiFeO3-0.3BaTiO3 Lead-free Piezoceramics
by: Liu, Jinming, et al.
Published: (2025)
by: Liu, Jinming, et al.
Published: (2025)
An Ensemble Framework for Explainable Geospatial Machine Learning Models
by: Liu, Lingbo
Published: (2024)
by: Liu, Lingbo
Published: (2024)
Learning When to See and When to Feel: Adaptive Vision-Torque Fusion for Contact-Aware Manipulation
by: Lei, Jiuzhou, et al.
Published: (2026)
by: Lei, Jiuzhou, et al.
Published: (2026)
Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation
by: Yao, Yuanqi, et al.
Published: (2025)
by: Yao, Yuanqi, et al.
Published: (2025)
Learning to Act Through Contact: A Unified View of Multi-Task Robot Learning
by: Omar, Shafeef, et al.
Published: (2025)
by: Omar, Shafeef, et al.
Published: (2025)
RoboView-Bias: Benchmarking Visual Bias in Embodied Agents for Robotic Manipulation
by: Liu, Enguang, et al.
Published: (2025)
by: Liu, Enguang, et al.
Published: (2025)
Contact SLAM: An Active Tactile Exploration Policy Based on Physical Reasoning Utilized in Robotic Fine Blind Manipulation Tasks
by: Wang, Gaozhao, et al.
Published: (2025)
by: Wang, Gaozhao, et al.
Published: (2025)
On the Complexity of Minimizing Energy Consumption of Partitioning DAG Tasks
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
EquAct: An SE(3)-Equivariant Multi-Task Transformer for Open-Loop Robotic Manipulation
by: Zhu, Xupeng, et al.
Published: (2025)
by: Zhu, Xupeng, et al.
Published: (2025)
Precise Object and Effect Removal with Adaptive Target-Aware Attention
by: Zhao, Jixin, et al.
Published: (2025)
by: Zhao, Jixin, et al.
Published: (2025)
Intrinsic Language-Guided Exploration for Complex Long-Horizon Robotic Manipulation Tasks
by: Triantafyllidis, Eleftherios, et al.
Published: (2023)
by: Triantafyllidis, Eleftherios, et al.
Published: (2023)
Ada3Drift: Adaptive Training-Time Drifting for One-Step 3D Visuomotor Robotic Manipulation
by: Xu, Chongyang, et al.
Published: (2026)
by: Xu, Chongyang, et al.
Published: (2026)
Act to See, See to Act: Diffusion-Driven Perception-Action Interplay for Adaptive Policies
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
FoAM: Foresight-Augmented Multi-Task Imitation Policy for Robotic Manipulation
by: Liu, Litao, et al.
Published: (2024)
by: Liu, Litao, et al.
Published: (2024)
Similar Items
-
RoVer: Robot Reward Model as Test-Time Verifier for Vision-Language-Action Model
by: Dai, Mingtong, et al.
Published: (2025) -
SkiP: When to Skip and When to Refine for Efficient Robot Manipulation
by: Dai, Mingtong, et al.
Published: (2026) -
Beyond the Destination: A Novel Benchmark for Exploration-Aware Embodied Question Answering
by: Jiang, Kaixuan, et al.
Published: (2025) -
GraspView: Active Perception Scoring and Best-View Optimization for Robotic Grasping in Cluttered Environments
by: Wang, Shenglin, et al.
Published: (2025) -
Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
by: Liu, Yang, et al.
Published: (2024)