Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Clark, Christopher, Zhang, Jieyu, Ma, Zixian, Park, Jae Sung, Salehi, Mohammadreza, Tripathi, Rohun, Lee, Sangho, Ren, Zhongzheng, Kim, Chris Dongjoo, Yang, Yinuo, Shao, Vincent, Yang, Yue, Huang, Weikai, Gao, Ziqi, Anderson, Taira, Zhang, Jianrui, Jain, Jitesh, Stoica, George, Han, Winson, Farhadi, Ali, Krishna, Ranjay |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
by: Clark, Christopher, et al.
Published: (2026)
by: Clark, Christopher, et al.
Published: (2026)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
by: Jain, Jitesh, et al.
Published: (2025)
by: Jain, Jitesh, et al.
Published: (2025)
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
by: Gupta, Tanmay, et al.
Published: (2026)
by: Gupta, Tanmay, et al.
Published: (2026)
You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass
by: Yang, Yinuo, et al.
Published: (2026)
by: Yang, Yinuo, et al.
Published: (2026)
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
by: Ma, Zixian, et al.
Published: (2024)
by: Ma, Zixian, et al.
Published: (2024)
MolmoAct: Action Reasoning Models that can Reason in Space
by: Lee, Jason, et al.
Published: (2025)
by: Lee, Jason, et al.
Published: (2025)
VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition
by: Yadav, Tanush, et al.
Published: (2026)
by: Yadav, Tanush, et al.
Published: (2026)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
by: Gao, Ziqi, et al.
Published: (2024)
by: Gao, Ziqi, et al.
Published: (2024)
Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
by: Huang, Weikai, et al.
Published: (2025)
by: Huang, Weikai, et al.
Published: (2025)
WildDet3D: Scaling Promptable 3D Detection in the Wild
by: Huang, Weikai, et al.
Published: (2026)
by: Huang, Weikai, et al.
Published: (2026)
MolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulation
by: Deshpande, Abhay, et al.
Published: (2026)
by: Deshpande, Abhay, et al.
Published: (2026)
One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory
by: Zheng, Chenhao, et al.
Published: (2025)
by: Zheng, Chenhao, et al.
Published: (2025)
Task Me Anything
by: Zhang, Jieyu, et al.
Published: (2024)
by: Zhang, Jieyu, et al.
Published: (2024)
MolmoAct2: Action Reasoning Models for Real-world Deployment
by: Fang, Haoquan, et al.
Published: (2026)
by: Fang, Haoquan, et al.
Published: (2026)
GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation
by: Deshpande, Abhay, et al.
Published: (2025)
by: Deshpande, Abhay, et al.
Published: (2025)
Synthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos
by: Gao, Ziqi, et al.
Published: (2026)
by: Gao, Ziqi, et al.
Published: (2026)
Posterior Augmented Flow Matching
by: Stoica, George, et al.
Published: (2026)
by: Stoica, George, et al.
Published: (2026)
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
by: Deitke, Matt, et al.
Published: (2024)
by: Deitke, Matt, et al.
Published: (2024)
Contrastive Flow Matching
by: Stoica, George, et al.
Published: (2025)
by: Stoica, George, et al.
Published: (2025)
MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
by: Kim, Yejin, et al.
Published: (2026)
by: Kim, Yejin, et al.
Published: (2026)
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
by: Salehi, Mohammadreza, et al.
Published: (2024)
by: Salehi, Mohammadreza, et al.
Published: (2024)
Iterated Learning Improves Compositionality in Large Vision-Language Models
by: Zheng, Chenhao, et al.
Published: (2024)
by: Zheng, Chenhao, et al.
Published: (2024)
Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
by: Li, Linjie, et al.
Published: (2025)
by: Li, Linjie, et al.
Published: (2025)
TrajTok: Learning Trajectory Tokens enables better Video Understanding
by: Zheng, Chenhao, et al.
Published: (2026)
by: Zheng, Chenhao, et al.
Published: (2026)
RefDecoder: Enhancing Visual Generation with Conditional Video Decoding
by: Fan, Xiang, et al.
Published: (2026)
by: Fan, Xiang, et al.
Published: (2026)
Visual Representations inside the Language Model
by: Liu, Benlin, et al.
Published: (2025)
by: Liu, Benlin, et al.
Published: (2025)
Video-Based Reward Modeling for Computer-Use Agents
by: Song, Linxin, et al.
Published: (2026)
by: Song, Linxin, et al.
Published: (2026)
ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models
by: Zhang, Jieyu, et al.
Published: (2024)
by: Zhang, Jieyu, et al.
Published: (2024)
Towards Acyclic Preference Evaluation of Language Models via Multiple Evaluators
by: Hu, Zhengyu, et al.
Published: (2024)
by: Hu, Zhengyu, et al.
Published: (2024)
One Diffusion to Generate Them All
by: Le, Duong H., et al.
Published: (2024)
by: Le, Duong H., et al.
Published: (2024)
Offline Training of Language Model Agents with Functions as Learnable Weights
by: Zhang, Shaokun, et al.
Published: (2024)
by: Zhang, Shaokun, et al.
Published: (2024)
vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models
by: Choi, Suhwan, et al.
Published: (2026)
by: Choi, Suhwan, et al.
Published: (2026)
Structure From Tracking: Distilling Structure-Preserving Motion for Video Generation
by: Fei, Yang, et al.
Published: (2025)
by: Fei, Yang, et al.
Published: (2025)
MutaGReP: Execution-Free Repository-Grounded Plan Search for Code-Use
by: Khan, Zaid, et al.
Published: (2025)
by: Khan, Zaid, et al.
Published: (2025)
Rethinking Human Preference Evaluation of LLM Rationales
by: Li, Ziang, et al.
Published: (2025)
by: Li, Ziang, et al.
Published: (2025)
Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
by: Jain, Jitesh, et al.
Published: (2024)
by: Jain, Jitesh, et al.
Published: (2024)
REALEDIT: Reddit Edits As a Large-scale Empirical Dataset for Image Transformations
by: Sushko, Peter, et al.
Published: (2025)
by: Sushko, Peter, et al.
Published: (2025)
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
by: Zhang, Tianyi, et al.
Published: (2026)
by: Zhang, Tianyi, et al.
Published: (2026)
LATTE: Learning to Think with Vision Specialists
by: Ma, Zixian, et al.
Published: (2024)
by: Ma, Zixian, et al.
Published: (2024)
Similar Items
-
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
by: Clark, Christopher, et al.
Published: (2026) -
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
by: Zhang, Jianrui, et al.
Published: (2026) -
SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
by: Jain, Jitesh, et al.
Published: (2025) -
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
by: Gupta, Tanmay, et al.
Published: (2026) -
You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass
by: Yang, Yinuo, et al.
Published: (2026)