Scaling Instructable Agents Across Many Simulated Worlds
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912068668489728 |
|---|---|
| author | SIMA Team Raad, Maria Abi Ahuja, Arun Barros, Catarina Besse, Frederic Bolt, Andrew Bolton, Adrian Brownfield, Bethanie Buttimore, Gavin Cant, Max Chakera, Sarah Chan, Stephanie C. Y. Clune, Jeff Collister, Adrian Copeman, Vikki Cullum, Alex Dasgupta, Ishita de Cesare, Dario Di Trapani, Julia Donchev, Yani Dunleavy, Emma Engelcke, Martin Faulkner, Ryan Garcia, Frankie Gbadamosi, Charles Gong, Zhitao Gonzales, Lucy Gupta, Kshitij Gregor, Karol Hallingstad, Arne Olav Harley, Tim Haves, Sam Hill, Felix Hirst, Ed Hudson, Drew A. Hudson, Jony Hughes-Fitt, Steph Rezende, Danilo J. Jasarevic, Mimi Kampis, Laura Ke, Rosemary Keck, Thomas Kim, Junkyung Knagg, Oscar Kopparapu, Kavya Lawton, Rory Lampinen, Andrew Legg, Shane Lerchner, Alexander Limont, Marjorie Liu, Yulan Loks-Thompson, Maria Marino, Joseph Cussons, Kathryn Martin Matthey, Loic Mcloughlin, Siobhan Mendolicchio, Piermaria Merzic, Hamza Mitenkova, Anna Moufarek, Alexandre Oliveira, Valeria Oliveira, Yanko Openshaw, Hannah Pan, Renke Pappu, Aneesh Platonov, Alex Purkiss, Ollie Reichert, David Reid, John Richemond, Pierre Harvey Roberts, Tyson Ruscoe, Giles Elias, Jaume Sanchez Sandars, Tasha Sawyer, Daniel P. Scholtes, Tim Simmons, Guy Slater, Daniel Soyer, Hubert Strathmann, Heiko Stys, Peter Tam, Allison C. Teplyashin, Denis Terzi, Tayfun Vercelli, Davide Vujatovic, Bojan Wainwright, Marcus Wang, Jane X. Wang, Zhengdong Wierstra, Daan Williams, Duncan Wong, Nathaniel York, Sarah Young, Nick |
| author_facet | SIMA Team Raad, Maria Abi Ahuja, Arun Barros, Catarina Besse, Frederic Bolt, Andrew Bolton, Adrian Brownfield, Bethanie Buttimore, Gavin Cant, Max Chakera, Sarah Chan, Stephanie C. Y. Clune, Jeff Collister, Adrian Copeman, Vikki Cullum, Alex Dasgupta, Ishita de Cesare, Dario Di Trapani, Julia Donchev, Yani Dunleavy, Emma Engelcke, Martin Faulkner, Ryan Garcia, Frankie Gbadamosi, Charles Gong, Zhitao Gonzales, Lucy Gupta, Kshitij Gregor, Karol Hallingstad, Arne Olav Harley, Tim Haves, Sam Hill, Felix Hirst, Ed Hudson, Drew A. Hudson, Jony Hughes-Fitt, Steph Rezende, Danilo J. Jasarevic, Mimi Kampis, Laura Ke, Rosemary Keck, Thomas Kim, Junkyung Knagg, Oscar Kopparapu, Kavya Lawton, Rory Lampinen, Andrew Legg, Shane Lerchner, Alexander Limont, Marjorie Liu, Yulan Loks-Thompson, Maria Marino, Joseph Cussons, Kathryn Martin Matthey, Loic Mcloughlin, Siobhan Mendolicchio, Piermaria Merzic, Hamza Mitenkova, Anna Moufarek, Alexandre Oliveira, Valeria Oliveira, Yanko Openshaw, Hannah Pan, Renke Pappu, Aneesh Platonov, Alex Purkiss, Ollie Reichert, David Reid, John Richemond, Pierre Harvey Roberts, Tyson Ruscoe, Giles Elias, Jaume Sanchez Sandars, Tasha Sawyer, Daniel P. Scholtes, Tim Simmons, Guy Slater, Daniel Soyer, Hubert Strathmann, Heiko Stys, Peter Tam, Allison C. Teplyashin, Denis Terzi, Tayfun Vercelli, Davide Vujatovic, Bojan Wainwright, Marcus Wang, Jane X. Wang, Zhengdong Wierstra, Daan Williams, Duncan Wong, Nathaniel York, Sarah Young, Nick |
| contents | Building embodied AI systems that can follow arbitrary language instructions in any 3D environment is a key challenge for creating general AI. Accomplishing this goal requires learning to ground language in perception and embodied actions, in order to accomplish complex tasks. The Scalable, Instructable, Multiworld Agent (SIMA) project tackles this by training agents to follow free-form instructions across a diverse range of virtual 3D environments, including curated research environments as well as open-ended, commercial video games. Our goal is to develop an instructable agent that can accomplish anything a human can do in any simulated 3D environment. Our approach focuses on language-driven generality while imposing minimal assumptions. Our agents interact with environments in real-time using a generic, human-like interface: the inputs are image observations and language instructions and the outputs are keyboard-and-mouse actions. This general approach is challenging, but it allows agents to ground language across many visually complex and semantically rich environments while also allowing us to readily run agents in new environments. In this paper we describe our motivation and goal, the initial progress we have made, and promising preliminary results on several diverse research environments and a variety of commercial video games. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2404_10179 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Scaling Instructable Agents Across Many Simulated Worlds SIMA Team Raad, Maria Abi Ahuja, Arun Barros, Catarina Besse, Frederic Bolt, Andrew Bolton, Adrian Brownfield, Bethanie Buttimore, Gavin Cant, Max Chakera, Sarah Chan, Stephanie C. Y. Clune, Jeff Collister, Adrian Copeman, Vikki Cullum, Alex Dasgupta, Ishita de Cesare, Dario Di Trapani, Julia Donchev, Yani Dunleavy, Emma Engelcke, Martin Faulkner, Ryan Garcia, Frankie Gbadamosi, Charles Gong, Zhitao Gonzales, Lucy Gupta, Kshitij Gregor, Karol Hallingstad, Arne Olav Harley, Tim Haves, Sam Hill, Felix Hirst, Ed Hudson, Drew A. Hudson, Jony Hughes-Fitt, Steph Rezende, Danilo J. Jasarevic, Mimi Kampis, Laura Ke, Rosemary Keck, Thomas Kim, Junkyung Knagg, Oscar Kopparapu, Kavya Lawton, Rory Lampinen, Andrew Legg, Shane Lerchner, Alexander Limont, Marjorie Liu, Yulan Loks-Thompson, Maria Marino, Joseph Cussons, Kathryn Martin Matthey, Loic Mcloughlin, Siobhan Mendolicchio, Piermaria Merzic, Hamza Mitenkova, Anna Moufarek, Alexandre Oliveira, Valeria Oliveira, Yanko Openshaw, Hannah Pan, Renke Pappu, Aneesh Platonov, Alex Purkiss, Ollie Reichert, David Reid, John Richemond, Pierre Harvey Roberts, Tyson Ruscoe, Giles Elias, Jaume Sanchez Sandars, Tasha Sawyer, Daniel P. Scholtes, Tim Simmons, Guy Slater, Daniel Soyer, Hubert Strathmann, Heiko Stys, Peter Tam, Allison C. Teplyashin, Denis Terzi, Tayfun Vercelli, Davide Vujatovic, Bojan Wainwright, Marcus Wang, Jane X. Wang, Zhengdong Wierstra, Daan Williams, Duncan Wong, Nathaniel York, Sarah Young, Nick Robotics Artificial Intelligence Human-Computer Interaction Machine Learning Building embodied AI systems that can follow arbitrary language instructions in any 3D environment is a key challenge for creating general AI. Accomplishing this goal requires learning to ground language in perception and embodied actions, in order to accomplish complex tasks. The Scalable, Instructable, Multiworld Agent (SIMA) project tackles this by training agents to follow free-form instructions across a diverse range of virtual 3D environments, including curated research environments as well as open-ended, commercial video games. Our goal is to develop an instructable agent that can accomplish anything a human can do in any simulated 3D environment. Our approach focuses on language-driven generality while imposing minimal assumptions. Our agents interact with environments in real-time using a generic, human-like interface: the inputs are image observations and language instructions and the outputs are keyboard-and-mouse actions. This general approach is challenging, but it allows agents to ground language across many visually complex and semantically rich environments while also allowing us to readily run agents in new environments. In this paper we describe our motivation and goal, the initial progress we have made, and promising preliminary results on several diverse research environments and a variety of commercial video games. |
| title | Scaling Instructable Agents Across Many Simulated Worlds |
| topic | Robotics Artificial Intelligence Human-Computer Interaction Machine Learning |
| url | https://arxiv.org/abs/2404.10179 |