MolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Deshpande, Abhay, Guru, Maya, Hendrix, Rose, Jauhri, Snehal, Eftekhar, Ainaz, Tripathi, Rohun, Argus, Max, Salvador, Jordi, Fang, Haoquan, Wallingford, Matthew, Pumacay, Wilbert, Kim, Yejin, Pfeifer, Quinn, Lee, Ying-Chun, Wolters, Piper, Rayyan, Omar, Zhang, Mingtong, Duan, Jiafei, Farley, Karen, Han, Winson, Vanderbilt, Eli, Fox, Dieter, Farhadi, Ali, Chalvatzaki, Georgia, Shah, Dhruv, Krishna, Ranjay |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
von: Kim, Yejin, et al.
Veröffentlicht: (2026)
von: Kim, Yejin, et al.
Veröffentlicht: (2026)
MolmoAct: Action Reasoning Models that can Reason in Space
von: Lee, Jason, et al.
Veröffentlicht: (2025)
von: Lee, Jason, et al.
Veröffentlicht: (2025)
Selective Visual Representations Improve Convergence and Generalization for Embodied AI
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2023)
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2023)
MolmoAct2: Action Reasoning Models for Real-world Deployment
von: Fang, Haoquan, et al.
Veröffentlicht: (2026)
von: Fang, Haoquan, et al.
Veröffentlicht: (2026)
SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation
von: Fang, Haoquan, et al.
Veröffentlicht: (2025)
von: Fang, Haoquan, et al.
Veröffentlicht: (2025)
Learning Any-View 6DoF Robotic Grasping in Cluttered Scenes via Neural Surface Rendering
von: Jauhri, Snehal, et al.
Veröffentlicht: (2023)
von: Jauhri, Snehal, et al.
Veröffentlicht: (2023)
Whole-Body Mobile Manipulation using Offline Reinforcement Learning on Sub-optimal Controllers
von: Jauhri, Snehal, et al.
Veröffentlicht: (2026)
von: Jauhri, Snehal, et al.
Veröffentlicht: (2026)
Active-Perceptive Motion Generation for Mobile Manipulation
von: Jauhri, Snehal, et al.
Veröffentlicht: (2023)
von: Jauhri, Snehal, et al.
Veröffentlicht: (2023)
GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation
von: Deshpande, Abhay, et al.
Veröffentlicht: (2025)
von: Deshpande, Abhay, et al.
Veröffentlicht: (2025)
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
von: Clark, Christopher, et al.
Veröffentlicht: (2026)
von: Clark, Christopher, et al.
Veröffentlicht: (2026)
THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation
von: Pumacay, Wilbert, et al.
Veröffentlicht: (2024)
von: Pumacay, Wilbert, et al.
Veröffentlicht: (2024)
UniFField: A Generalizable Unified Neural Feature Field for Visual, Semantic, and Spatial Uncertainties in Any Scene
von: Maurer, Christian, et al.
Veröffentlicht: (2025)
von: Maurer, Christian, et al.
Veröffentlicht: (2025)
2HandedAfforder: Learning Precise Actionable Bimanual Affordances from Human Videos
von: Heidinger, Marvin, et al.
Veröffentlicht: (2025)
von: Heidinger, Marvin, et al.
Veröffentlicht: (2025)
Convergent Functions, Divergent Forms
von: Jeon, Hyeonseong, et al.
Veröffentlicht: (2025)
von: Jeon, Hyeonseong, et al.
Veröffentlicht: (2025)
The One RING: a Robotic Indoor Navigation Generalist
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2024)
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2024)
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
von: Gupta, Tanmay, et al.
Veröffentlicht: (2026)
von: Gupta, Tanmay, et al.
Veröffentlicht: (2026)
6DOPE-GS: Online 6D Object Pose Estimation using Gaussian Splatting
von: Jin, Yufeng, et al.
Veröffentlicht: (2024)
von: Jin, Yufeng, et al.
Veröffentlicht: (2024)
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
von: Yuan, Wentao, et al.
Veröffentlicht: (2024)
von: Yuan, Wentao, et al.
Veröffentlicht: (2024)
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
von: Clark, Christopher, et al.
Veröffentlicht: (2026)
von: Clark, Christopher, et al.
Veröffentlicht: (2026)
Posterior Augmented Flow Matching
von: Stoica, George, et al.
Veröffentlicht: (2026)
von: Stoica, George, et al.
Veröffentlicht: (2026)
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
von: Cheng, Long, et al.
Veröffentlicht: (2025)
von: Cheng, Long, et al.
Veröffentlicht: (2025)
Think Twice, Act Once: Verifier-Guided Action Selection For Embodied Agents
von: Singhi, Nishad, et al.
Veröffentlicht: (2026)
von: Singhi, Nishad, et al.
Veröffentlicht: (2026)
AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition
von: Yadav, Tanush, et al.
Veröffentlicht: (2026)
von: Yadav, Tanush, et al.
Veröffentlicht: (2026)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
von: Zhang, Jianrui, et al.
Veröffentlicht: (2026)
von: Zhang, Jianrui, et al.
Veröffentlicht: (2026)
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
von: Deitke, Matt, et al.
Veröffentlicht: (2024)
von: Deitke, Matt, et al.
Veröffentlicht: (2024)
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
von: Lin, Zijun, et al.
Veröffentlicht: (2025)
von: Lin, Zijun, et al.
Veröffentlicht: (2025)
Recurrent-Depth VLA: Implicit Test-Time Compute Scaling of Vision-Language-Action Models via Latent Iterative Reasoning
von: Tur, Yalcin, et al.
Veröffentlicht: (2026)
von: Tur, Yalcin, et al.
Veröffentlicht: (2026)
WildDet3D: Scaling Promptable 3D Detection in the Wild
von: Huang, Weikai, et al.
Veröffentlicht: (2026)
von: Huang, Weikai, et al.
Veröffentlicht: (2026)
RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation
von: Wang, Yi Ru, et al.
Veröffentlicht: (2025)
von: Wang, Yi Ru, et al.
Veröffentlicht: (2025)
Effect of Pullulan‐Chitosan Film Containing Titanium Dioxide (TiO 2 ) Nanoparticle and Tarragon Essential Oil on the Quality Properties During Refrigerated Storage of Rainbow Trout Fillets
von: Ainaz Khodanazary
Veröffentlicht: (2025)
von: Ainaz Khodanazary
Veröffentlicht: (2025)
Effects of Carboxymethyl Chitosan/Pectin Coating Containing Free and Nanoliposome Mentha piperita Essential Oil on the Shelf Life of Shrimp During Ice Storage
von: Ainaz Khodanazary
Veröffentlicht: (2025)
von: Ainaz Khodanazary
Veröffentlicht: (2025)
Beyond the Frame: Generating 360 Panoramic Videos from Perspective Videos
von: Luo, Rundong, et al.
Veröffentlicht: (2025)
von: Luo, Rundong, et al.
Veröffentlicht: (2025)
Sustaining multidisciplinary teams in rural and remote primary care
von: Geoff Argus
Veröffentlicht: (2024)
von: Geoff Argus
Veröffentlicht: (2024)
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
von: Ahmad, Ghazi Shazan, et al.
Veröffentlicht: (2025)
von: Ahmad, Ghazi Shazan, et al.
Veröffentlicht: (2025)
1979 annual survey of American law / Vanderbilt Hall, Arthur T
von: Vanderbilt Hall, Arthur T
Veröffentlicht: (1979)
von: Vanderbilt Hall, Arthur T
Veröffentlicht: (1979)
Connecting Learning: Brain-Based Strategies for Linking Prior Knowledge in the Library Media Center
von: Vanderbilt, Kathi L.
Veröffentlicht: (2005)
von: Vanderbilt, Kathi L.
Veröffentlicht: (2005)
Superposed Decoding: Multiple Generations from a Single Autoregressive Inference Pass
von: Shen, Ethan, et al.
Veröffentlicht: (2024)
von: Shen, Ethan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
von: Kim, Yejin, et al.
Veröffentlicht: (2026) -
MolmoAct: Action Reasoning Models that can Reason in Space
von: Lee, Jason, et al.
Veröffentlicht: (2025) -
Selective Visual Representations Improve Convergence and Generalization for Embodied AI
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2023) -
MolmoAct2: Action Reasoning Models for Real-world Deployment
von: Fang, Haoquan, et al.
Veröffentlicht: (2026) -
SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation
von: Fang, Haoquan, et al.
Veröffentlicht: (2025)