Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Deitke, Matt, Clark, Christopher, Lee, Sangho, Tripathi, Rohun, Yang, Yue, Park, Jae Sung, Salehi, Mohammadreza, Muennighoff, Niklas, Lo, Kyle, Soldaini, Luca, Lu, Jiasen, Anderson, Taira, Bransom, Erin, Ehsani, Kiana, Ngo, Huong, Chen, YenSung, Patel, Ajay, Yatskar, Mark, Callison-Burch, Chris, Head, Andrew, Hendrix, Rose, Bastani, Favyen, VanderBilt, Eli, Lambert, Nathan, Chou, Yvonne, Chheda, Arnavi, Sparks, Jenna, Skjonsberg, Sam, Schmitz, Michael, Sarnat, Aaron, Bischoff, Byron, Walsh, Pete, Newell, Chris, Wolters, Piper, Gupta, Tanmay, Zeng, Kuo-Hao, Borchardt, Jon, Groeneveld, Dirk, Nam, Crystal, Lebrecht, Sophie, Wittlif, Caitlin, Schoenick, Carissa, Michel, Oscar, Krishna, Ranjay, Weihs, Luca, Smith, Noah A., Hajishirzi, Hannaneh, Girshick, Ross, Farhadi, Ali, Kembhavi, Aniruddha |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
di: Clark, Christopher, et al.
Pubblicazione: (2026)
di: Clark, Christopher, et al.
Pubblicazione: (2026)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
di: Yang, Yue, et al.
Pubblicazione: (2025)
di: Yang, Yue, et al.
Pubblicazione: (2025)
MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
di: Kim, Yejin, et al.
Pubblicazione: (2026)
di: Kim, Yejin, et al.
Pubblicazione: (2026)
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
di: Clark, Christopher, et al.
Pubblicazione: (2026)
di: Clark, Christopher, et al.
Pubblicazione: (2026)
OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens
di: Liu, Jiacheng, et al.
Pubblicazione: (2025)
di: Liu, Jiacheng, et al.
Pubblicazione: (2025)
Holodeck: Language Guided Generation of 3D Embodied AI Environments
di: Yang, Yue, et al.
Pubblicazione: (2023)
di: Yang, Yue, et al.
Pubblicazione: (2023)
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
di: Gupta, Tanmay, et al.
Pubblicazione: (2026)
di: Gupta, Tanmay, et al.
Pubblicazione: (2026)
Engaging with Children's Artwork in Mixed Visual-Ability Families
di: Chheda-Kothary, Arnavi, et al.
Pubblicazione: (2024)
di: Chheda-Kothary, Arnavi, et al.
Pubblicazione: (2024)
MolmoAct: Action Reasoning Models that can Reason in Space
di: Lee, Jason, et al.
Pubblicazione: (2025)
di: Lee, Jason, et al.
Pubblicazione: (2025)
MolmoAct2: Action Reasoning Models for Real-world Deployment
di: Fang, Haoquan, et al.
Pubblicazione: (2026)
di: Fang, Haoquan, et al.
Pubblicazione: (2026)
ArtInsight: Enabling AI-Powered Artwork Engagement for Mixed Visual-Ability Families
di: Chheda-Kothary, Arnavi, et al.
Pubblicazione: (2025)
di: Chheda-Kothary, Arnavi, et al.
Pubblicazione: (2025)
Postmodernism in Youth Literature--A Road Away from the Reader?
di: Skjonsberg, Kari
Pubblicazione: (1992)
di: Skjonsberg, Kari
Pubblicazione: (1992)
PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators
di: Zeng, Kuo-Hao, et al.
Pubblicazione: (2024)
di: Zeng, Kuo-Hao, et al.
Pubblicazione: (2024)
CodeNav: Beyond tool-use to using real-world codebases with LLM agents
di: Gupta, Tanmay, et al.
Pubblicazione: (2024)
di: Gupta, Tanmay, et al.
Pubblicazione: (2024)
AltGeoViz: Facilitating Accessible Geovisualization
di: Li, Chu, et al.
Pubblicazione: (2024)
di: Li, Chu, et al.
Pubblicazione: (2024)
Open Educational Resources: An Annotated Bibliography for Librarians
di: Bober, Chris
Pubblicazione: (2017)
di: Bober, Chris
Pubblicazione: (2017)
SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World
di: Ehsani, Kiana, et al.
Pubblicazione: (2023)
di: Ehsani, Kiana, et al.
Pubblicazione: (2023)
olmOCR 2: Unit Test Rewards for Document OCR
di: Poznanski, Jake, et al.
Pubblicazione: (2025)
di: Poznanski, Jake, et al.
Pubblicazione: (2025)
OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
di: Ngo, Huong, et al.
Pubblicazione: (2025)
di: Ngo, Huong, et al.
Pubblicazione: (2025)
MolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulation
di: Deshpande, Abhay, et al.
Pubblicazione: (2026)
di: Deshpande, Abhay, et al.
Pubblicazione: (2026)
Diagrama de fase sitio-enlace en dímeros
di: W. Lebrecht
Pubblicazione: (2014)
di: W. Lebrecht
Pubblicazione: (2014)
Umbrales de percolación de sitios. Pequeñas celdas bidimensionales asimétricas
di: W. Lebrecht
Pubblicazione: (2009)
di: W. Lebrecht
Pubblicazione: (2009)
Frustración en redes arquimedianas antiferromagnéticas
di: W. Lebrecht
Pubblicazione: (2013)
di: W. Lebrecht
Pubblicazione: (2013)
Frustración local en la red arquimediana (3,4,6,4)
di: W. Lebrecht
Pubblicazione: (2008)
di: W. Lebrecht
Pubblicazione: (2008)
Umbrales de percolación exactos en redes duales
di: W. Lebrecht
Pubblicazione: (2010)
di: W. Lebrecht
Pubblicazione: (2010)
Percolación discreta en redes tridimensionales
di: W. Lebrecht
Pubblicazione: (2011)
di: W. Lebrecht
Pubblicazione: (2011)
Umbral de percolación en las redes de Kagomé y Dice
di: W. Lebrecht
Pubblicazione: (2013)
di: W. Lebrecht
Pubblicazione: (2013)
GeoVisA11y: An AI-based Geovisualization Question-Answering System for Screen-Reader Users
di: Li, Chu, et al.
Pubblicazione: (2026)
di: Li, Chu, et al.
Pubblicazione: (2026)
Structuring data analysis projects in the Open Science era with Kerblam!
di: Visentin, Luca, et al.
Pubblicazione: (2024)
di: Visentin, Luca, et al.
Pubblicazione: (2024)
How2Everything: Mining the Web for How-To Procedures to Evaluate and Improve LLMs
di: Chang, Yapei, et al.
Pubblicazione: (2026)
di: Chang, Yapei, et al.
Pubblicazione: (2026)
Ensemble Transformer for Efficient and Accurate Ranking Tasks: an Application to Question Answering Systems
di: Matsubara, Yoshitomo, et al.
Pubblicazione: (2022)
di: Matsubara, Yoshitomo, et al.
Pubblicazione: (2022)
The Path to Open Innovation: Peer-Review Under Fire (PRUF)
di: Billions, Ava, et al.
Pubblicazione: (2025)
di: Billions, Ava, et al.
Pubblicazione: (2025)
Taming Offload Overheads in a Massively Parallel Open-Source RISC-V MPSoC: Analysis and Optimization
di: Colagrande, Luca, et al.
Pubblicazione: (2025)
di: Colagrande, Luca, et al.
Pubblicazione: (2025)
OpenSlot: Mixed Open-Set Recognition with Object-Centric Learning
di: Yin, Xu, et al.
Pubblicazione: (2024)
di: Yin, Xu, et al.
Pubblicazione: (2024)
Amuse: Human-AI Collaborative Songwriting with Multimodal Inspirations
di: Kim, Yewon, et al.
Pubblicazione: (2024)
di: Kim, Yewon, et al.
Pubblicazione: (2024)
TOP: Towards Open & Predictable Heterogeneous SoCs
di: Valente, Luca, et al.
Pubblicazione: (2024)
di: Valente, Luca, et al.
Pubblicazione: (2024)
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
di: Ahmad, Ghazi Shazan, et al.
Pubblicazione: (2025)
di: Ahmad, Ghazi Shazan, et al.
Pubblicazione: (2025)
OLMoE: Open Mixture-of-Experts Language Models
di: Muennighoff, Niklas, et al.
Pubblicazione: (2024)
di: Muennighoff, Niklas, et al.
Pubblicazione: (2024)
FlexOlmo: Open Language Models for Flexible Data Use
di: Shi, Weijia, et al.
Pubblicazione: (2025)
di: Shi, Weijia, et al.
Pubblicazione: (2025)
Bayesian Fields: Task-driven Open-Set Semantic Gaussian Splatting
di: Maggio, Dominic, et al.
Pubblicazione: (2025)
di: Maggio, Dominic, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
di: Clark, Christopher, et al.
Pubblicazione: (2026) -
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
di: Yang, Yue, et al.
Pubblicazione: (2025) -
MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
di: Kim, Yejin, et al.
Pubblicazione: (2026) -
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
di: Clark, Christopher, et al.
Pubblicazione: (2026) -
OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens
di: Liu, Jiacheng, et al.
Pubblicazione: (2025)