Act to See, See to Act: Diffusion-Driven Perception-Action Interplay for Adaptive Policies
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Jing, Peng, Weiting, Tang, Jing, Gong, Zeyu, Wang, Xihua, Tao, Bo, Cheng, Li |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
par: Xu, Kechun, et autres
Publié: (2025)
par: Xu, Kechun, et autres
Publié: (2025)
See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
par: Chen, Guangyan, et autres
Publié: (2025)
par: Chen, Guangyan, et autres
Publié: (2025)
Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
par: Bai, Yongjie, et autres
Publié: (2025)
par: Bai, Yongjie, et autres
Publié: (2025)
MemoAct: Atkinson-Shiffrin-Inspired Memory-Augmented Visuomotor Policy for Robotic Manipulation
par: Tan, Liufan, et autres
Publié: (2026)
par: Tan, Liufan, et autres
Publié: (2026)
Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation
par: Zuo, Kuangji, et autres
Publié: (2026)
par: Zuo, Kuangji, et autres
Publié: (2026)
Act, Sense, Act: Learning Non-Markovian Active Perception Strategies from Large-Scale Egocentric Human Data
par: Li, Jialiang, et autres
Publié: (2026)
par: Li, Jialiang, et autres
Publié: (2026)
ThermoAct:Thermal-Aware Vision-Language-Action Models for Robotic Perception and Decision-Making
par: Son, Young-Chae, et autres
Publié: (2026)
par: Son, Young-Chae, et autres
Publié: (2026)
MolmoAct: Action Reasoning Models that can Reason in Space
par: Lee, Jason, et autres
Publié: (2025)
par: Lee, Jason, et autres
Publié: (2025)
MolmoAct2: Action Reasoning Models for Real-world Deployment
par: Fang, Haoquan, et autres
Publié: (2026)
par: Fang, Haoquan, et autres
Publié: (2026)
See Tomorrow, Act Today: Foresight-Driven Autonomous Driving
par: Zhang, Bozhou, et autres
Publié: (2026)
par: Zhang, Bozhou, et autres
Publié: (2026)
Eye, Robot: Learning to Look to Act with a BC-RL Perception-Action Loop
par: Kerr, Justin, et autres
Publié: (2025)
par: Kerr, Justin, et autres
Publié: (2025)
Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation
par: Wang, Guokang, et autres
Publié: (2024)
par: Wang, Guokang, et autres
Publié: (2024)
VisualActBench: Can VLMs See and Act like a Human?
par: Zhang, Daoan, et autres
Publié: (2025)
par: Zhang, Daoan, et autres
Publié: (2025)
UniAct: Unified Motion Generation and Action Streaming for Humanoid Robots
par: Jiang, Nan, et autres
Publié: (2025)
par: Jiang, Nan, et autres
Publié: (2025)
AdaWorldPolicy: World-Model-Driven Diffusion Policy with Online Adaptive Learning for Robotic Manipulation
par: Yuan, Ge, et autres
Publié: (2026)
par: Yuan, Ge, et autres
Publié: (2026)
Safe-Night VLA: Seeing the Unseen via Thermal-Perceptive Vision-Language-Action Models for Safety-Critical Manipulation
par: Yu, Dian, et autres
Publié: (2026)
par: Yu, Dian, et autres
Publié: (2026)
Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation
par: Hui, Chenyu, et autres
Publié: (2026)
par: Hui, Chenyu, et autres
Publié: (2026)
RoboAct-CLIP: Video-Driven Pre-training of Atomic Action Understanding for Robotics
par: Zhang, Zhiyuan, et autres
Publié: (2025)
par: Zhang, Zhiyuan, et autres
Publié: (2025)
Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models
par: Izzo, Riccardo Andrea, et autres
Publié: (2026)
par: Izzo, Riccardo Andrea, et autres
Publié: (2026)
Say, Dream, and Act: Learning Video World Models for Instruction-Driven Robot Manipulation
par: Gu, Songen, et autres
Publié: (2026)
par: Gu, Songen, et autres
Publié: (2026)
FlowAct: A Proactive Multimodal Human-robot Interaction System with Continuous Flow of Perception and Modular Action Sub-systems
par: Dhaussy, Timothée, et autres
Publié: (2024)
par: Dhaussy, Timothée, et autres
Publié: (2024)
Imagine2Act: Leveraging Object-Action Motion Consistency from Imagined Goals for Robotic Manipulation
par: Heng, Liang, et autres
Publié: (2025)
par: Heng, Liang, et autres
Publié: (2025)
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
par: Ye, Wencheng, et autres
Publié: (2025)
par: Ye, Wencheng, et autres
Publié: (2025)
ForceFlow: Learning to Feel and Act via Contact-Driven Flow Matching
par: Zhang, Shuoheng, et autres
Publié: (2026)
par: Zhang, Shuoheng, et autres
Publié: (2026)
Learning to Act Robustly with View-Invariant Latent Actions
par: Jeong, Youngjoon, et autres
Publié: (2026)
par: Jeong, Youngjoon, et autres
Publié: (2026)
Grasp, See, and Place: Efficient Unknown Object Rearrangement with Policy Structure Prior
par: Xu, Kechun, et autres
Publié: (2024)
par: Xu, Kechun, et autres
Publié: (2024)
When to Act, Ask, or Learn: Uncertainty-Aware Policy Steering
par: Yuan, Jessie, et autres
Publié: (2026)
par: Yuan, Jessie, et autres
Publié: (2026)
UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
par: Bu, Qingwen, et autres
Publié: (2025)
par: Bu, Qingwen, et autres
Publié: (2025)
Seeing Beyond: Extrapolative Domain Adaptive Panoramic Segmentation
par: Zheng, Yuanfan, et autres
Publié: (2026)
par: Zheng, Yuanfan, et autres
Publié: (2026)
VANP: Learning Where to See for Navigation with Self-Supervised Vision-Action Pre-Training
par: Nazeri, Mohammad, et autres
Publié: (2024)
par: Nazeri, Mohammad, et autres
Publié: (2024)
See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
par: Zhang, Yimeng, et autres
Publié: (2025)
par: Zhang, Yimeng, et autres
Publié: (2025)
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
par: Chi, Cheng, et autres
Publié: (2023)
par: Chi, Cheng, et autres
Publié: (2023)
Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling
par: Jia, Yueru, et autres
Publié: (2025)
par: Jia, Yueru, et autres
Publié: (2025)
Seeing before Observable: Potential Risk Reasoning in Autonomous Driving via Vision Language Models
par: Liu, Jiaxin, et autres
Publié: (2025)
par: Liu, Jiaxin, et autres
Publié: (2025)
Seeing the Bigger Picture: 3D Latent Mapping for Mobile Manipulation Policy Learning
par: Kim, Sunghwan, et autres
Publié: (2025)
par: Kim, Sunghwan, et autres
Publié: (2025)
StreamVLA: Breaking the Reason-Act Cycle via Completion-State Gating
par: Chen, Tongqing, et autres
Publié: (2026)
par: Chen, Tongqing, et autres
Publié: (2026)
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models
par: Ling, Yiran, et autres
Publié: (2026)
par: Ling, Yiran, et autres
Publié: (2026)
Act2See: Emergent Active Visual Perception for Video Reasoning
par: Ma, Martin Q., et autres
Publié: (2026)
par: Ma, Martin Q., et autres
Publié: (2026)
Microscopic Robots That Sense, Think, Act, and Compute
par: Lassiter, Maya M., et autres
Publié: (2025)
par: Lassiter, Maya M., et autres
Publié: (2025)
Select before Act: Spatially Decoupled Action Repetition for Continuous Control
par: Nie, Buqing, et autres
Publié: (2025)
par: Nie, Buqing, et autres
Publié: (2025)
Documents similaires
-
Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
par: Xu, Kechun, et autres
Publié: (2025) -
See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
par: Chen, Guangyan, et autres
Publié: (2025) -
Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
par: Bai, Yongjie, et autres
Publié: (2025) -
MemoAct: Atkinson-Shiffrin-Inspired Memory-Augmented Visuomotor Policy for Robotic Manipulation
par: Tan, Liufan, et autres
Publié: (2026) -
Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation
par: Zuo, Kuangji, et autres
Publié: (2026)