_version_ 1866917002757537792
author FAIR CodeGen team
Copet, Jade
Carbonneaux, Quentin
Cohen, Gal
Gehring, Jonas
Kahn, Jacob
Kossen, Jannik
Kreuk, Felix
McMilin, Emily
Meyer, Michel
Wei, Yuxiang
Zhang, David
Zheng, Kunhao
Armengol-Estapé, Jordi
Bashiri, Pedram
Beck, Maximilian
Chambon, Pierre
Charnalia, Abhishek
Cummins, Chris
Decugis, Juliette
Fisches, Zacharias V.
Fleuret, François
Gloeckle, Fabian
Gu, Alex
Hassid, Michael
Haziza, Daniel
Idrissi, Badr Youbi
Keller, Christian
Kindi, Rahul
Leather, Hugh
Maimon, Gallil
Markosyan, Aram
Massa, Francisco
Mazaré, Pierre-Emmanuel
Mella, Vegard
Murray, Naila
Muzumdar, Keyur
O'Hearn, Peter
Pagliardini, Matteo
Pedchenko, Dmitrii
Remez, Tal
Seeker, Volker
Selvi, Marco
Sultan, Oren
Wang, Sida
Wehrstedt, Luca
Yoran, Ori
Zhang, Lingming
Cohen, Taco
Adi, Yossi
Synnaeve, Gabriel
author_facet FAIR CodeGen team
Copet, Jade
Carbonneaux, Quentin
Cohen, Gal
Gehring, Jonas
Kahn, Jacob
Kossen, Jannik
Kreuk, Felix
McMilin, Emily
Meyer, Michel
Wei, Yuxiang
Zhang, David
Zheng, Kunhao
Armengol-Estapé, Jordi
Bashiri, Pedram
Beck, Maximilian
Chambon, Pierre
Charnalia, Abhishek
Cummins, Chris
Decugis, Juliette
Fisches, Zacharias V.
Fleuret, François
Gloeckle, Fabian
Gu, Alex
Hassid, Michael
Haziza, Daniel
Idrissi, Badr Youbi
Keller, Christian
Kindi, Rahul
Leather, Hugh
Maimon, Gallil
Markosyan, Aram
Massa, Francisco
Mazaré, Pierre-Emmanuel
Mella, Vegard
Murray, Naila
Muzumdar, Keyur
O'Hearn, Peter
Pagliardini, Matteo
Pedchenko, Dmitrii
Remez, Tal
Seeker, Volker
Selvi, Marco
Sultan, Oren
Wang, Sida
Wehrstedt, Luca
Yoran, Ori
Zhang, Lingming
Cohen, Taco
Adi, Yossi
Synnaeve, Gabriel
contents We release Code World Model (CWM), a 32-billion-parameter open-weights LLM, to advance research on code generation with world models. To improve code understanding beyond what can be learned from training on static code alone, we mid-train CWM on a large amount of observation-action trajectories from Python interpreter and agentic Docker environments, and perform extensive multi-task reasoning RL in verifiable coding, math, and multi-turn software engineering environments. With CWM, we provide a strong testbed for researchers to explore the opportunities world modeling affords for improving code generation with reasoning and planning in computational environments. We present first steps of how world models can benefit agentic coding, enable step-by-step simulation of Python code execution, and show early results of how reasoning can benefit from the latter. CWM is a dense, decoder-only LLM trained with a context size of up to 131k tokens. Independent of its world modeling capabilities, CWM offers strong performance on general coding and math tasks: it reaches pass@1 scores of 65.8% on SWE-bench Verified (with test-time scaling), 68.6% on LiveCodeBench, 96.6% on Math-500, and 76.0% on AIME 2024. To support further research on code world modeling, we release model checkpoints after mid-training, SFT, and RL.
format Preprint
id arxiv_https___arxiv_org_abs_2510_02387
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CWM: An Open-Weights LLM for Research on Code Generation with World Models
FAIR CodeGen team
Copet, Jade
Carbonneaux, Quentin
Cohen, Gal
Gehring, Jonas
Kahn, Jacob
Kossen, Jannik
Kreuk, Felix
McMilin, Emily
Meyer, Michel
Wei, Yuxiang
Zhang, David
Zheng, Kunhao
Armengol-Estapé, Jordi
Bashiri, Pedram
Beck, Maximilian
Chambon, Pierre
Charnalia, Abhishek
Cummins, Chris
Decugis, Juliette
Fisches, Zacharias V.
Fleuret, François
Gloeckle, Fabian
Gu, Alex
Hassid, Michael
Haziza, Daniel
Idrissi, Badr Youbi
Keller, Christian
Kindi, Rahul
Leather, Hugh
Maimon, Gallil
Markosyan, Aram
Massa, Francisco
Mazaré, Pierre-Emmanuel
Mella, Vegard
Murray, Naila
Muzumdar, Keyur
O'Hearn, Peter
Pagliardini, Matteo
Pedchenko, Dmitrii
Remez, Tal
Seeker, Volker
Selvi, Marco
Sultan, Oren
Wang, Sida
Wehrstedt, Luca
Yoran, Ori
Zhang, Lingming
Cohen, Taco
Adi, Yossi
Synnaeve, Gabriel
Software Engineering
Artificial Intelligence
Machine Learning
68T07
I.2.7
We release Code World Model (CWM), a 32-billion-parameter open-weights LLM, to advance research on code generation with world models. To improve code understanding beyond what can be learned from training on static code alone, we mid-train CWM on a large amount of observation-action trajectories from Python interpreter and agentic Docker environments, and perform extensive multi-task reasoning RL in verifiable coding, math, and multi-turn software engineering environments. With CWM, we provide a strong testbed for researchers to explore the opportunities world modeling affords for improving code generation with reasoning and planning in computational environments. We present first steps of how world models can benefit agentic coding, enable step-by-step simulation of Python code execution, and show early results of how reasoning can benefit from the latter. CWM is a dense, decoder-only LLM trained with a context size of up to 131k tokens. Independent of its world modeling capabilities, CWM offers strong performance on general coding and math tasks: it reaches pass@1 scores of 65.8% on SWE-bench Verified (with test-time scaling), 68.6% on LiveCodeBench, 96.6% on Math-500, and 76.0% on AIME 2024. To support further research on code world modeling, we release model checkpoints after mid-training, SFT, and RL.
title CWM: An Open-Weights LLM for Research on Code Generation with World Models
topic Software Engineering
Artificial Intelligence
Machine Learning
68T07
I.2.7
url https://arxiv.org/abs/2510.02387