MolmoAct: Action Reasoning Models that can Reason in Space
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Jason, Duan, Jiafei, Fang, Haoquan, Deng, Yuquan, Liu, Shuo, Li, Boyang, Fang, Bohan, Zhang, Jieyu, Wang, Yi Ru, Lee, Sangho, Han, Winson, Pumacay, Wilbert, Wu, Angelica, Hendrix, Rose, Farley, Karen, VanderBilt, Eli, Farhadi, Ali, Fox, Dieter, Krishna, Ranjay |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MolmoAct2: Action Reasoning Models for Real-world Deployment
by: Fang, Haoquan, et al.
Published: (2026)
by: Fang, Haoquan, et al.
Published: (2026)
SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation
by: Fang, Haoquan, et al.
Published: (2025)
by: Fang, Haoquan, et al.
Published: (2025)
MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
by: Kim, Yejin, et al.
Published: (2026)
by: Kim, Yejin, et al.
Published: (2026)
MolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulation
by: Deshpande, Abhay, et al.
Published: (2026)
by: Deshpande, Abhay, et al.
Published: (2026)
GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation
by: Deshpande, Abhay, et al.
Published: (2025)
by: Deshpande, Abhay, et al.
Published: (2025)
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
by: Clark, Christopher, et al.
Published: (2026)
by: Clark, Christopher, et al.
Published: (2026)
THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation
by: Pumacay, Wilbert, et al.
Published: (2024)
by: Pumacay, Wilbert, et al.
Published: (2024)
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
by: Lin, Zijun, et al.
Published: (2025)
by: Lin, Zijun, et al.
Published: (2025)
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
by: Duan, Jiafei, et al.
Published: (2024)
by: Duan, Jiafei, et al.
Published: (2024)
AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
by: Duan, Jiafei, et al.
Published: (2024)
by: Duan, Jiafei, et al.
Published: (2024)
Recurrent-Depth VLA: Implicit Test-Time Compute Scaling of Vision-Language-Action Models via Latent Iterative Reasoning
by: Tur, Yalcin, et al.
Published: (2026)
by: Tur, Yalcin, et al.
Published: (2026)
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
by: Yuan, Wentao, et al.
Published: (2024)
by: Yuan, Wentao, et al.
Published: (2024)
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
by: Clark, Christopher, et al.
Published: (2026)
by: Clark, Christopher, et al.
Published: (2026)
SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World
by: Ehsani, Kiana, et al.
Published: (2023)
by: Ehsani, Kiana, et al.
Published: (2023)
The One RING: a Robotic Indoor Navigation Generalist
by: Eftekhar, Ainaz, et al.
Published: (2024)
by: Eftekhar, Ainaz, et al.
Published: (2024)
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
by: Gupta, Tanmay, et al.
Published: (2026)
by: Gupta, Tanmay, et al.
Published: (2026)
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
by: Deitke, Matt, et al.
Published: (2024)
by: Deitke, Matt, et al.
Published: (2024)
RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation
by: Wang, Yi Ru, et al.
Published: (2025)
by: Wang, Yi Ru, et al.
Published: (2025)
Holodeck: Language Guided Generation of 3D Embodied AI Environments
by: Yang, Yue, et al.
Published: (2023)
by: Yang, Yue, et al.
Published: (2023)
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
by: Cheng, Long, et al.
Published: (2025)
by: Cheng, Long, et al.
Published: (2025)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
Selective Visual Representations Improve Convergence and Generalization for Embodied AI
by: Eftekhar, Ainaz, et al.
Published: (2023)
by: Eftekhar, Ainaz, et al.
Published: (2023)
EVE: Enabling Anyone to Train Robots using Augmented Reality
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
by: Jain, Jitesh, et al.
Published: (2025)
by: Jain, Jitesh, et al.
Published: (2025)
RefDecoder: Enhancing Visual Generation with Conditional Video Decoding
by: Fan, Xiang, et al.
Published: (2026)
by: Fan, Xiang, et al.
Published: (2026)
Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
by: Huang, Weikai, et al.
Published: (2025)
by: Huang, Weikai, et al.
Published: (2025)
Posterior Augmented Flow Matching
by: Stoica, George, et al.
Published: (2026)
by: Stoica, George, et al.
Published: (2026)
WildDet3D: Scaling Promptable 3D Detection in the Wild
by: Huang, Weikai, et al.
Published: (2026)
by: Huang, Weikai, et al.
Published: (2026)
StreamVLA: Breaking the Reason-Act Cycle via Completion-State Gating
by: Chen, Tongqing, et al.
Published: (2026)
by: Chen, Tongqing, et al.
Published: (2026)
Optimize Cardinality Estimation Model Pretraining by Simplifying the Training Datasets
by: Fang, Boyang
Published: (2025)
by: Fang, Boyang
Published: (2025)
Self-Contradictory Reasoning Evaluation and Detection
by: Liu, Ziyi, et al.
Published: (2023)
by: Liu, Ziyi, et al.
Published: (2023)
TEEMATE: Fast and Efficient Confidential Container using Shared Enclave
by: Lee, Chulmin, et al.
Published: (2024)
by: Lee, Chulmin, et al.
Published: (2024)
GraphReAct: Reasoning and Acting for Multi-step Graph Inference
by: Yu, Xingtong, et al.
Published: (2026)
by: Yu, Xingtong, et al.
Published: (2026)
Iterated Learning Improves Compositionality in Large Vision-Language Models
by: Zheng, Chenhao, et al.
Published: (2024)
by: Zheng, Chenhao, et al.
Published: (2024)
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
by: Ahmad, Ghazi Shazan, et al.
Published: (2025)
by: Ahmad, Ghazi Shazan, et al.
Published: (2025)
Visual Representations inside the Language Model
by: Liu, Benlin, et al.
Published: (2025)
by: Liu, Benlin, et al.
Published: (2025)
Self-DC: When to Reason and When to Act? Self Divide-and-Conquer for Compositional Unknown Questions
by: Wang, Hongru, et al.
Published: (2024)
by: Wang, Hongru, et al.
Published: (2024)
TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics
by: Chen, Shirui, et al.
Published: (2026)
by: Chen, Shirui, et al.
Published: (2026)
Can Language Models Laugh at YouTube Short-form Videos?
by: Ko, Dayoon, et al.
Published: (2023)
by: Ko, Dayoon, et al.
Published: (2023)
Allergic contact dermatitis to edible essential oils: A case report
by: Sangho Lee, et al.
Published: (2024)
by: Sangho Lee, et al.
Published: (2024)
Similar Items
-
MolmoAct2: Action Reasoning Models for Real-world Deployment
by: Fang, Haoquan, et al.
Published: (2026) -
SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation
by: Fang, Haoquan, et al.
Published: (2025) -
MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
by: Kim, Yejin, et al.
Published: (2026) -
MolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulation
by: Deshpande, Abhay, et al.
Published: (2026) -
GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation
by: Deshpande, Abhay, et al.
Published: (2025)