MolmoAct2: Action Reasoning Models for Real-world Deployment
Fuente:
arXiv
Saved in:
| Main Authors: | Fang, Haoquan, Duan, Jiafei, Clay, Donovan, Wang, Sam, Liu, Shuo, Huang, Weikai, Fan, Xiang, Tsai, Wei-Chuan, Chen, Shirui, Wang, Yi Ru, Xing, Shanli, Cho, Jaemin, Park, Jae Sung, Eftekhar, Ainaz, Sushko, Peter, Farley, Karen, Wadhwa, Angad, Harrison, Cole, Han, Winson, Lee, Ying-Chun, VanderBilt, Eli, Hendrix, Rose, Ellawela, Suveen, Ngoo, Lucas, Chai, Joyce, Ren, Zhongzheng, Farhadi, Ali, Fox, Dieter, Krishna, Ranjay |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MolmoAct: Action Reasoning Models that can Reason in Space
by: Lee, Jason, et al.
Published: (2025)
by: Lee, Jason, et al.
Published: (2025)
MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
by: Kim, Yejin, et al.
Published: (2026)
by: Kim, Yejin, et al.
Published: (2026)
Trust, Lies, and Long Memories: Emergent Social Dynamics and Reputation in Multi-Round Avalon with LLM Agents
by: Ellawela, Suveen
Published: (2026)
by: Ellawela, Suveen
Published: (2026)
MolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulation
by: Deshpande, Abhay, et al.
Published: (2026)
by: Deshpande, Abhay, et al.
Published: (2026)
Selective Visual Representations Improve Convergence and Generalization for Embodied AI
by: Eftekhar, Ainaz, et al.
Published: (2023)
by: Eftekhar, Ainaz, et al.
Published: (2023)
GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation
by: Deshpande, Abhay, et al.
Published: (2025)
by: Deshpande, Abhay, et al.
Published: (2025)
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
by: Gupta, Tanmay, et al.
Published: (2026)
by: Gupta, Tanmay, et al.
Published: (2026)
The One RING: a Robotic Indoor Navigation Generalist
by: Eftekhar, Ainaz, et al.
Published: (2024)
by: Eftekhar, Ainaz, et al.
Published: (2024)
Proactive Agentic Whiteboards: Enhancing Diagrammatic Learning
by: Ellawela, Suveen, et al.
Published: (2025)
by: Ellawela, Suveen, et al.
Published: (2025)
Convergent Functions, Divergent Forms
by: Jeon, Hyeonseong, et al.
Published: (2025)
by: Jeon, Hyeonseong, et al.
Published: (2025)
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
by: Clark, Christopher, et al.
Published: (2026)
by: Clark, Christopher, et al.
Published: (2026)
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
by: Clark, Christopher, et al.
Published: (2026)
by: Clark, Christopher, et al.
Published: (2026)
SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World
by: Ehsani, Kiana, et al.
Published: (2023)
by: Ehsani, Kiana, et al.
Published: (2023)
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
by: Cheng, Long, et al.
Published: (2025)
by: Cheng, Long, et al.
Published: (2025)
WildDet3D: Scaling Promptable 3D Detection in the Wild
by: Huang, Weikai, et al.
Published: (2026)
by: Huang, Weikai, et al.
Published: (2026)
Holodeck: Language Guided Generation of 3D Embodied AI Environments
by: Yang, Yue, et al.
Published: (2023)
by: Yang, Yue, et al.
Published: (2023)
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
by: Deitke, Matt, et al.
Published: (2024)
by: Deitke, Matt, et al.
Published: (2024)
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
by: Lin, Zijun, et al.
Published: (2025)
by: Lin, Zijun, et al.
Published: (2025)
TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics
by: Chen, Shirui, et al.
Published: (2026)
by: Chen, Shirui, et al.
Published: (2026)
SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation
by: Fang, Haoquan, et al.
Published: (2025)
by: Fang, Haoquan, et al.
Published: (2025)
Recurrent-Depth VLA: Implicit Test-Time Compute Scaling of Vision-Language-Action Models via Latent Iterative Reasoning
by: Tur, Yalcin, et al.
Published: (2026)
by: Tur, Yalcin, et al.
Published: (2026)
Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
by: Huang, Weikai, et al.
Published: (2025)
by: Huang, Weikai, et al.
Published: (2025)
Posterior Augmented Flow Matching
by: Stoica, George, et al.
Published: (2026)
by: Stoica, George, et al.
Published: (2026)
Effect of Pullulan‐Chitosan Film Containing Titanium Dioxide (TiO 2 ) Nanoparticle and Tarragon Essential Oil on the Quality Properties During Refrigerated Storage of Rainbow Trout Fillets
by: Ainaz Khodanazary
Published: (2025)
by: Ainaz Khodanazary
Published: (2025)
Effects of Carboxymethyl Chitosan/Pectin Coating Containing Free and Nanoliposome Mentha piperita Essential Oil on the Shelf Life of Shrimp During Ice Storage
by: Ainaz Khodanazary
Published: (2025)
by: Ainaz Khodanazary
Published: (2025)
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
by: Ahmad, Ghazi Shazan, et al.
Published: (2025)
by: Ahmad, Ghazi Shazan, et al.
Published: (2025)
RefDecoder: Enhancing Visual Generation with Conditional Video Decoding
by: Fan, Xiang, et al.
Published: (2026)
by: Fan, Xiang, et al.
Published: (2026)
The impact of constipation management on the quality of life of Alzheimer’s patients
by: Viacheslav Viktorovich Sushko, et al.
Published: (2024)
by: Viacheslav Viktorovich Sushko, et al.
Published: (2024)
Melatonin to treat sleep disturbances in patients with Alzheimer’s disease
by: Viacheslav Viktorovich Sushko, et al.
Published: (2024)
by: Viacheslav Viktorovich Sushko, et al.
Published: (2024)
Visual Representations inside the Language Model
by: Liu, Benlin, et al.
Published: (2025)
by: Liu, Benlin, et al.
Published: (2025)
Comprehensive Analysis of the Full Tess Orbital Phase Curve of Wasp-121b
by: M. Eftekhar
Published: (2022)
by: M. Eftekhar
Published: (2022)
Exploring expectations of Iranian audiences in terms of consecutive interpreting: A reception study
by: Elnaz Eftekhar
Published: (2024)
by: Elnaz Eftekhar
Published: (2024)
Adversarial Latent-State Training for Robust Policies in Partially Observable Domains
by: Ahuja, Angad Singh
Published: (2026)
by: Ahuja, Angad Singh
Published: (2026)
Generalized Hilbert Operator Acting on Hardy Spaces
by: Chen, Huiling, et al.
Published: (2024)
by: Chen, Huiling, et al.
Published: (2024)
A Derivative-Hilbert operator acting on BMOA space
by: Chen, Huiling, et al.
Published: (2024)
by: Chen, Huiling, et al.
Published: (2024)
Generalized Hilbert operators acting from Hardy spaces to weighted Bergman spaces
by: Wang, Liyi, et al.
Published: (2025)
by: Wang, Liyi, et al.
Published: (2025)
Norm of the Hilbert matrix operator on logarithmically weighted Bloch and Hardy spaces
by: Ye, Shanli, et al.
Published: (2025)
by: Ye, Shanli, et al.
Published: (2025)
Norm of the Hilbert matrix operator between some spaces of analytic functions
by: Hu, Hao, et al.
Published: (2024)
by: Hu, Hao, et al.
Published: (2024)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
by: Gao, Ziqi, et al.
Published: (2024)
by: Gao, Ziqi, et al.
Published: (2024)
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
by: Ma, Zixian, et al.
Published: (2024)
by: Ma, Zixian, et al.
Published: (2024)
Similar Items
-
MolmoAct: Action Reasoning Models that can Reason in Space
by: Lee, Jason, et al.
Published: (2025) -
MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
by: Kim, Yejin, et al.
Published: (2026) -
Trust, Lies, and Long Memories: Emergent Social Dynamics and Reputation in Multi-Round Avalon with LLM Agents
by: Ellawela, Suveen
Published: (2026) -
MolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulation
by: Deshpande, Abhay, et al.
Published: (2026) -
Selective Visual Representations Improve Convergence and Generalization for Embodied AI
by: Eftekhar, Ainaz, et al.
Published: (2023)