Act, Sense, Act: Learning Non-Markovian Active Perception Strategies from Large-Scale Egocentric Human Data
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jialiang, Qiao, Yi, Guo, Yunhan, Chen, Changwen, Lian, Wenzhao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ActMVS: Active Scene Reconstruction with Monocular Multi-View Stereo
by: Pu, Guo, et al.
Published: (2026)
by: Pu, Guo, et al.
Published: (2026)
BridgeACT: Bridging Human Demonstrations to Robot Actions via Unified Tool-Target Affordances
by: Han, Yifan, et al.
Published: (2026)
by: Han, Yifan, et al.
Published: (2026)
Act to See, See to Act: Diffusion-Driven Perception-Action Interplay for Adaptive Policies
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
Microscopic Robots That Sense, Think, Act, and Compute
by: Lassiter, Maya M., et al.
Published: (2025)
by: Lassiter, Maya M., et al.
Published: (2025)
Behavior Cloning for Active Perception with Low-Resolution Egocentric Vision
by: Bilic, Anthony, et al.
Published: (2026)
by: Bilic, Anthony, et al.
Published: (2026)
ActSafe: Active Exploration with Safety Constraints for Reinforcement Learning
by: As, Yarden, et al.
Published: (2024)
by: As, Yarden, et al.
Published: (2024)
SAGE: Scene Graph-Aware Guidance and Execution for Long-Horizon Manipulation Tasks
by: Li, Jialiang, et al.
Published: (2025)
by: Li, Jialiang, et al.
Published: (2025)
Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation
by: Wang, Guokang, et al.
Published: (2024)
by: Wang, Guokang, et al.
Published: (2024)
EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data
by: Zheng, Ruijie, et al.
Published: (2026)
by: Zheng, Ruijie, et al.
Published: (2026)
EMMA: Scaling Mobile Manipulation via Egocentric Human Data
by: Zhu, Lawrence Y., et al.
Published: (2025)
by: Zhu, Lawrence Y., et al.
Published: (2025)
Eye, Robot: Learning to Look to Act with a BC-RL Perception-Action Loop
by: Kerr, Justin, et al.
Published: (2025)
by: Kerr, Justin, et al.
Published: (2025)
ActLoc: Learning to Localize on the Move via Active Viewpoint Selection
by: Li, Jiajie, et al.
Published: (2025)
by: Li, Jiajie, et al.
Published: (2025)
TesserAct: Learning 4D Embodied World Models
by: Zhen, Haoyu, et al.
Published: (2025)
by: Zhen, Haoyu, et al.
Published: (2025)
Look, Zoom, Understand: The Robotic Eyeball for Embodied Perception
by: Yang, Jiashu, et al.
Published: (2025)
by: Yang, Jiashu, et al.
Published: (2025)
EgoLive: A Large-Scale Egocentric Dataset from Real-World Human Tasks
by: Li, Yihang, et al.
Published: (2026)
by: Li, Yihang, et al.
Published: (2026)
FlowAct: A Proactive Multimodal Human-robot Interaction System with Continuous Flow of Perception and Modular Action Sub-systems
by: Dhaussy, Timothée, et al.
Published: (2024)
by: Dhaussy, Timothée, et al.
Published: (2024)
Retrieval-Augmented Robots via Retrieve-Reason-Act
by: Temiraliev, Izat, et al.
Published: (2026)
by: Temiraliev, Izat, et al.
Published: (2026)
ThermoAct:Thermal-Aware Vision-Language-Action Models for Robotic Perception and Decision-Making
by: Son, Young-Chae, et al.
Published: (2026)
by: Son, Young-Chae, et al.
Published: (2026)
Act Better by Timing: A timing-Aware Reinforcement Learning for Autonomous Driving
by: Li, Guanzhou, et al.
Published: (2024)
by: Li, Guanzhou, et al.
Published: (2024)
CLASH: Collision Learning via Augmented Sim-to-real Hybridization to Bridge the Reality Gap
by: He, Haotian, et al.
Published: (2026)
by: He, Haotian, et al.
Published: (2026)
EgoMI: Learning Active Vision and Whole-Body Manipulation from Egocentric Human Demonstrations
by: Yu, Justin, et al.
Published: (2025)
by: Yu, Justin, et al.
Published: (2025)
Prepare Before You Act: Learning From Humans to Rearrange Initial States
by: Dai, Yinlong, et al.
Published: (2025)
by: Dai, Yinlong, et al.
Published: (2025)
Plan-and-Act using Large Language Models for Interactive Agreement
by: Sasabuchi, Kazuhiro, et al.
Published: (2025)
by: Sasabuchi, Kazuhiro, et al.
Published: (2025)
COMPASS: Confined-space Manipulation Planning with Active Sensing Strategy
by: Li, Qixuan, et al.
Published: (2025)
by: Li, Qixuan, et al.
Published: (2025)
Non-Overlap-Aware Egocentric Pose Estimation for Collaborative Perception in Connected Autonomy
by: Huang, Hong, et al.
Published: (2025)
by: Huang, Hong, et al.
Published: (2025)
Instruct Large Language Models to Drive like Humans
by: Zhang, Ruijun, et al.
Published: (2024)
by: Zhang, Ruijun, et al.
Published: (2024)
MolmoAct: Action Reasoning Models that can Reason in Space
by: Lee, Jason, et al.
Published: (2025)
by: Lee, Jason, et al.
Published: (2025)
Viewpoint Matters: Dynamically Optimizing Viewpoints with Masked Autoencoder for Visual Manipulation
by: Yi, Pengfei, et al.
Published: (2026)
by: Yi, Pengfei, et al.
Published: (2026)
StreamVLA: Breaking the Reason-Act Cycle via Completion-State Gating
by: Chen, Tongqing, et al.
Published: (2026)
by: Chen, Tongqing, et al.
Published: (2026)
See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
by: Chen, Guangyan, et al.
Published: (2025)
by: Chen, Guangyan, et al.
Published: (2025)
PhysBrain: Human Egocentric Data as a Bridge from Vision Language Models to Physical Intelligence
by: Lin, Xiaopeng, et al.
Published: (2025)
by: Lin, Xiaopeng, et al.
Published: (2025)
Teaching Robots Where To Go And How To Act With Human Sketches via Spatial Diagrammatic Instructions
by: Sun, Qilin, et al.
Published: (2024)
by: Sun, Qilin, et al.
Published: (2024)
When to Act: Calibrated Confidence for Reliable Human Intention Prediction in Assistive Robotics
by: Gaus, Johannes A., et al.
Published: (2026)
by: Gaus, Johannes A., et al.
Published: (2026)
Learning to Act Through Contact: A Unified View of Multi-Task Robot Learning
by: Omar, Shafeef, et al.
Published: (2025)
by: Omar, Shafeef, et al.
Published: (2025)
When to Act, Ask, or Learn: Uncertainty-Aware Policy Steering
by: Yuan, Jessie, et al.
Published: (2026)
by: Yuan, Jessie, et al.
Published: (2026)
Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation?
by: Zhang, Zhongru, et al.
Published: (2026)
by: Zhang, Zhongru, et al.
Published: (2026)
Vision in Action: Learning Active Perception from Human Demonstrations
by: Xiong, Haoyu, et al.
Published: (2025)
by: Xiong, Haoyu, et al.
Published: (2025)
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models
by: Ling, Yiran, et al.
Published: (2026)
by: Ling, Yiran, et al.
Published: (2026)
Learning to Act Robustly with View-Invariant Latent Actions
by: Jeong, Youngjoon, et al.
Published: (2026)
by: Jeong, Youngjoon, et al.
Published: (2026)
MolmoAct2: Action Reasoning Models for Real-world Deployment
by: Fang, Haoquan, et al.
Published: (2026)
by: Fang, Haoquan, et al.
Published: (2026)
Similar Items
-
ActMVS: Active Scene Reconstruction with Monocular Multi-View Stereo
by: Pu, Guo, et al.
Published: (2026) -
BridgeACT: Bridging Human Demonstrations to Robot Actions via Unified Tool-Target Affordances
by: Han, Yifan, et al.
Published: (2026) -
Act to See, See to Act: Diffusion-Driven Perception-Action Interplay for Adaptive Policies
by: Wang, Jing, et al.
Published: (2025) -
Microscopic Robots That Sense, Think, Act, and Compute
by: Lassiter, Maya M., et al.
Published: (2025) -
Behavior Cloning for Active Perception with Low-Resolution Egocentric Vision
by: Bilic, Anthony, et al.
Published: (2026)