Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ganai, Milan, Luo, Katie, Frey, Jonas, Barrett, Clark, Pavone, Marco |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
von: NVIDIA, et al.
Veröffentlicht: (2025)
von: NVIDIA, et al.
Veröffentlicht: (2025)
RoaD: Rollouts as Demonstrations for Closed-Loop Supervised Fine-Tuning of Autonomous Driving Policies
von: Garcia-Cobo, Guillermo, et al.
Veröffentlicht: (2025)
von: Garcia-Cobo, Guillermo, et al.
Veröffentlicht: (2025)
Explore until Confident: Efficient Exploration for Embodied Question Answering
von: Ren, Allen Z., et al.
Veröffentlicht: (2024)
von: Ren, Allen Z., et al.
Veröffentlicht: (2024)
GrabS: Generative Embodied Agent for 3D Object Segmentation without Scene Supervision
von: Zhang, Zihui, et al.
Veröffentlicht: (2025)
von: Zhang, Zihui, et al.
Veröffentlicht: (2025)
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
von: Assran, Mido, et al.
Veröffentlicht: (2025)
von: Assran, Mido, et al.
Veröffentlicht: (2025)
Provable Ordering and Continuity in Vision-Language Pretraining for Generalizable Embodied Agents
von: Zhang, Zhizhen, et al.
Veröffentlicht: (2025)
von: Zhang, Zhizhen, et al.
Veröffentlicht: (2025)
ICAT: Incident-Case-Grounded Adaptive Testing for Physical-Risk Prediction in Embodied World Models
von: Lai, Zhenglin, et al.
Veröffentlicht: (2026)
von: Lai, Zhenglin, et al.
Veröffentlicht: (2026)
Grounding Video Models to Actions through Goal Conditioned Exploration
von: Luo, Yunhao, et al.
Veröffentlicht: (2024)
von: Luo, Yunhao, et al.
Veröffentlicht: (2024)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
GroCo: Ground Constraint for Metric Self-Supervised Monocular Depth
von: Cecille, Aurélien, et al.
Veröffentlicht: (2024)
von: Cecille, Aurélien, et al.
Veröffentlicht: (2024)
mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
von: Pai, Jonas, et al.
Veröffentlicht: (2025)
von: Pai, Jonas, et al.
Veröffentlicht: (2025)
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
von: Huang, Chi-Pin, et al.
Veröffentlicht: (2025)
von: Huang, Chi-Pin, et al.
Veröffentlicht: (2025)
GrndCtrl: Grounding World Models via Self-Supervised Reward Alignment
von: He, Haoyang, et al.
Veröffentlicht: (2025)
von: He, Haoyang, et al.
Veröffentlicht: (2025)
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
von: Hong, Yining, et al.
Veröffentlicht: (2026)
von: Hong, Yining, et al.
Veröffentlicht: (2026)
StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement
von: Seo, Junwon, et al.
Veröffentlicht: (2026)
von: Seo, Junwon, et al.
Veröffentlicht: (2026)
Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning
von: Huang, Chi-Pin, et al.
Veröffentlicht: (2026)
von: Huang, Chi-Pin, et al.
Veröffentlicht: (2026)
AutoWorld: Scaling Multi-Agent Traffic Simulation with Self-Supervised World Models
von: Pourkeshavatz, Mozhgan, et al.
Veröffentlicht: (2026)
von: Pourkeshavatz, Mozhgan, et al.
Veröffentlicht: (2026)
Vidar: Embodied Video Diffusion Model for Generalist Manipulation
von: Feng, Yao, et al.
Veröffentlicht: (2025)
von: Feng, Yao, et al.
Veröffentlicht: (2025)
EmbodiSwap for Zero-Shot Robot Imitation Learning
von: Dessalene, Eadom, et al.
Veröffentlicht: (2025)
von: Dessalene, Eadom, et al.
Veröffentlicht: (2025)
A Survey of Embodied Learning for Object-Centric Robotic Manipulation
von: Zheng, Ying, et al.
Veröffentlicht: (2024)
von: Zheng, Ying, et al.
Veröffentlicht: (2024)
Octopus: Embodied Vision-Language Programmer from Environmental Feedback
von: Yang, Jingkang, et al.
Veröffentlicht: (2023)
von: Yang, Jingkang, et al.
Veröffentlicht: (2023)
RoboFactory: Exploring Embodied Agent Collaboration with Compositional Constraints
von: Qin, Yiran, et al.
Veröffentlicht: (2025)
von: Qin, Yiran, et al.
Veröffentlicht: (2025)
Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?
von: Majumdar, Arjun, et al.
Veröffentlicht: (2023)
von: Majumdar, Arjun, et al.
Veröffentlicht: (2023)
RoboLayout: Differentiable 3D Scene Generation for Embodied Agents
von: Shamsaddinlou, Ali
Veröffentlicht: (2026)
von: Shamsaddinlou, Ali
Veröffentlicht: (2026)
NoisyEQA: Benchmarking Embodied Question Answering Against Noisy Queries
von: Wu, Tao, et al.
Veröffentlicht: (2024)
von: Wu, Tao, et al.
Veröffentlicht: (2024)
SPA: 3D Spatial-Awareness Enables Effective Embodied Representation
von: Zhu, Haoyi, et al.
Veröffentlicht: (2024)
von: Zhu, Haoyi, et al.
Veröffentlicht: (2024)
EfficientFlow: Efficient Equivariant Flow Policy Learning for Embodied AI
von: Chang, Jianlei, et al.
Veröffentlicht: (2025)
von: Chang, Jianlei, et al.
Veröffentlicht: (2025)
MonoPP: Metric-Scaled Self-Supervised Monocular Depth Estimation by Planar-Parallax Geometry in Automotive Applications
von: Elazab, Gasser, et al.
Veröffentlicht: (2024)
von: Elazab, Gasser, et al.
Veröffentlicht: (2024)
Extrapolated Urban View Synthesis Benchmark
von: Han, Xiangyu, et al.
Veröffentlicht: (2024)
von: Han, Xiangyu, et al.
Veröffentlicht: (2024)
HASARD: A Benchmark for Vision-Based Safe Reinforcement Learning in Embodied Agents
von: Tomilin, Tristan, et al.
Veröffentlicht: (2025)
von: Tomilin, Tristan, et al.
Veröffentlicht: (2025)
PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI
von: Yang, Yandan, et al.
Veröffentlicht: (2024)
von: Yang, Yandan, et al.
Veröffentlicht: (2024)
Wild Visual Navigation: Fast Traversability Learning via Pre-Trained Models and Online Self-Supervision
von: Mattamala, Matías, et al.
Veröffentlicht: (2024)
von: Mattamala, Matías, et al.
Veröffentlicht: (2024)
Pseudo-Simulation for Autonomous Driving
von: Cao, Wei, et al.
Veröffentlicht: (2025)
von: Cao, Wei, et al.
Veröffentlicht: (2025)
NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking
von: Dauner, Daniel, et al.
Veröffentlicht: (2024)
von: Dauner, Daniel, et al.
Veröffentlicht: (2024)
When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering
von: Wu, Tao, et al.
Veröffentlicht: (2025)
von: Wu, Tao, et al.
Veröffentlicht: (2025)
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
von: Zhang, Zhengshen, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengshen, et al.
Veröffentlicht: (2025)
3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding
von: Ogunleye, Makanjuola, et al.
Veröffentlicht: (2026)
von: Ogunleye, Makanjuola, et al.
Veröffentlicht: (2026)
SENSE: Self-Supervised Neural Embeddings for Spatial Ensembles
von: Gadirov, Hamid, et al.
Veröffentlicht: (2025)
von: Gadirov, Hamid, et al.
Veröffentlicht: (2025)
Learning to Visually Connect Actions and their Effects
von: Parmar, Paritosh, et al.
Veröffentlicht: (2024)
von: Parmar, Paritosh, et al.
Veröffentlicht: (2024)
EmboTeam: Grounding LLM Reasoning into Reactive Behavior Trees via PDDL for Embodied Multi-Robot Collaboration
von: Zeng, Haishan, et al.
Veröffentlicht: (2026)
von: Zeng, Haishan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
von: NVIDIA, et al.
Veröffentlicht: (2025) -
RoaD: Rollouts as Demonstrations for Closed-Loop Supervised Fine-Tuning of Autonomous Driving Policies
von: Garcia-Cobo, Guillermo, et al.
Veröffentlicht: (2025) -
Explore until Confident: Efficient Exploration for Embodied Question Answering
von: Ren, Allen Z., et al.
Veröffentlicht: (2024) -
GrabS: Generative Embodied Agent for 3D Object Segmentation without Scene Supervision
von: Zhang, Zihui, et al.
Veröffentlicht: (2025) -
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
von: Assran, Mido, et al.
Veröffentlicht: (2025)