The Open Molecules 2025 (OMol25) Dataset, Evaluations, and Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Levine, Daniel S., Shuaibi, Muhammed, Spotte-Smith, Evan Walter Clark, Taylor, Michael G., Hasyim, Muhammad R., Michel, Kyle, Batatia, Ilyes, Csányi, Gábor, Dzamba, Misko, Eastman, Peter, Frey, Nathan C., Fu, Xiang, Gharakhanyan, Vahe, Krishnapriyan, Aditi S., Rackers, Joshua A., Raja, Sanjeev, Rizvi, Ammar, Rosen, Andrew S., Ulissi, Zachary, Vargas, Santiago, Zitnick, C. Lawrence, Blau, Samuel M., Wood, Brandon M.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912942316847104
author Levine, Daniel S.
Shuaibi, Muhammed
Spotte-Smith, Evan Walter Clark
Taylor, Michael G.
Hasyim, Muhammad R.
Michel, Kyle
Batatia, Ilyes
Csányi, Gábor
Dzamba, Misko
Eastman, Peter
Frey, Nathan C.
Fu, Xiang
Gharakhanyan, Vahe
Krishnapriyan, Aditi S.
Rackers, Joshua A.
Raja, Sanjeev
Rizvi, Ammar
Rosen, Andrew S.
Ulissi, Zachary
Vargas, Santiago
Zitnick, C. Lawrence
Blau, Samuel M.
Wood, Brandon M.
author_facet Levine, Daniel S.
Shuaibi, Muhammed
Spotte-Smith, Evan Walter Clark
Taylor, Michael G.
Hasyim, Muhammad R.
Michel, Kyle
Batatia, Ilyes
Csányi, Gábor
Dzamba, Misko
Eastman, Peter
Frey, Nathan C.
Fu, Xiang
Gharakhanyan, Vahe
Krishnapriyan, Aditi S.
Rackers, Joshua A.
Raja, Sanjeev
Rizvi, Ammar
Rosen, Andrew S.
Ulissi, Zachary
Vargas, Santiago
Zitnick, C. Lawrence
Blau, Samuel M.
Wood, Brandon M.
contents Machine learning (ML) models hold the promise of transforming atomic simulations by delivering quantum chemical accuracy at a fraction of the computational cost. Realization of this potential would enable high-throughout, high-accuracy molecular screening campaigns to explore vast regions of chemical space and facilitate ab initio simulations at sizes and time scales that were previously inaccessible. However, a fundamental challenge to creating ML models that perform well across molecular chemistry is the lack of comprehensive data for training. Despite substantial efforts in data generation, no large-scale molecular dataset exists that combines broad chemical diversity with a high level of accuracy. To address this gap, Meta FAIR introduces Open Molecules 2025 (OMol25), a large-scale dataset composed of more than 100 million density functional theory (DFT) calculations at the $ω$B97M-V/def2-TZVPD level of theory, representing billions of CPU core-hours of compute. OMol25 uniquely blends elemental, chemical, and structural diversity including: 83 elements, a wide-range of intra- and intermolecular interactions, explicit solvation, variable charge/spin, conformers, and reactive structures. There are ~83M unique molecular systems in OMol25 covering small molecules, biomolecules, metal complexes, and electrolytes, including structures obtained from existing datasets. OMol25 also greatly expands on the size of systems typically included in DFT datasets, with systems of up to 350 atoms. In addition to the public release of the data, we provide baseline models and a comprehensive set of model evaluations to encourage community engagement in developing the next-generation ML models for molecular chemistry.
format Preprint
id arxiv_https___arxiv_org_abs_2505_08762
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Open Molecules 2025 (OMol25) Dataset, Evaluations, and Models
Levine, Daniel S.
Shuaibi, Muhammed
Spotte-Smith, Evan Walter Clark
Taylor, Michael G.
Hasyim, Muhammad R.
Michel, Kyle
Batatia, Ilyes
Csányi, Gábor
Dzamba, Misko
Eastman, Peter
Frey, Nathan C.
Fu, Xiang
Gharakhanyan, Vahe
Krishnapriyan, Aditi S.
Rackers, Joshua A.
Raja, Sanjeev
Rizvi, Ammar
Rosen, Andrew S.
Ulissi, Zachary
Vargas, Santiago
Zitnick, C. Lawrence
Blau, Samuel M.
Wood, Brandon M.
Chemical Physics
Machine learning (ML) models hold the promise of transforming atomic simulations by delivering quantum chemical accuracy at a fraction of the computational cost. Realization of this potential would enable high-throughout, high-accuracy molecular screening campaigns to explore vast regions of chemical space and facilitate ab initio simulations at sizes and time scales that were previously inaccessible. However, a fundamental challenge to creating ML models that perform well across molecular chemistry is the lack of comprehensive data for training. Despite substantial efforts in data generation, no large-scale molecular dataset exists that combines broad chemical diversity with a high level of accuracy. To address this gap, Meta FAIR introduces Open Molecules 2025 (OMol25), a large-scale dataset composed of more than 100 million density functional theory (DFT) calculations at the $ω$B97M-V/def2-TZVPD level of theory, representing billions of CPU core-hours of compute. OMol25 uniquely blends elemental, chemical, and structural diversity including: 83 elements, a wide-range of intra- and intermolecular interactions, explicit solvation, variable charge/spin, conformers, and reactive structures. There are ~83M unique molecular systems in OMol25 covering small molecules, biomolecules, metal complexes, and electrolytes, including structures obtained from existing datasets. OMol25 also greatly expands on the size of systems typically included in DFT datasets, with systems of up to 350 atoms. In addition to the public release of the data, we provide baseline models and a comprehensive set of model evaluations to encourage community engagement in developing the next-generation ML models for molecular chemistry.
title The Open Molecules 2025 (OMol25) Dataset, Evaluations, and Models
topic Chemical Physics
url https://arxiv.org/abs/2505.08762