An Interactive Agent Foundation Model

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Durante, Zane, Sarkar, Bidipta, Gong, Ran, Taori, Rohan, Noda, Yusuke, Tang, Paul, Adeli, Ehsan, Lakshmikanth, Shrinidhi Kowshika, Schulman, Kevin, Milstein, Arnold, Terzopoulos, Demetri, Famoti, Ade, Kuno, Noboru, Llorens, Ashley, Vo, Hoi, Ikeuchi, Katsu, Fei-Fei, Li, Gao, Jianfeng, Wake, Naoki, Huang, Qiuyuan
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910489840189440
author Durante, Zane
Sarkar, Bidipta
Gong, Ran
Taori, Rohan
Noda, Yusuke
Tang, Paul
Adeli, Ehsan
Lakshmikanth, Shrinidhi Kowshika
Schulman, Kevin
Milstein, Arnold
Terzopoulos, Demetri
Famoti, Ade
Kuno, Noboru
Llorens, Ashley
Vo, Hoi
Ikeuchi, Katsu
Fei-Fei, Li
Gao, Jianfeng
Wake, Naoki
Huang, Qiuyuan
author_facet Durante, Zane
Sarkar, Bidipta
Gong, Ran
Taori, Rohan
Noda, Yusuke
Tang, Paul
Adeli, Ehsan
Lakshmikanth, Shrinidhi Kowshika
Schulman, Kevin
Milstein, Arnold
Terzopoulos, Demetri
Famoti, Ade
Kuno, Noboru
Llorens, Ashley
Vo, Hoi
Ikeuchi, Katsu
Fei-Fei, Li
Gao, Jianfeng
Wake, Naoki
Huang, Qiuyuan
contents The development of artificial intelligence systems is transitioning from creating static, task-specific models to dynamic, agent-based systems capable of performing well in a wide range of applications. We propose an Interactive Agent Foundation Model that uses a novel multi-task agent training paradigm for training AI agents across a wide range of domains, datasets, and tasks. Our training paradigm unifies diverse pre-training strategies, including visual masked auto-encoders, language modeling, and next-action prediction, enabling a versatile and adaptable AI framework. We demonstrate the performance of our framework across three separate domains -- Robotics, Gaming AI, and Healthcare. Our model demonstrates its ability to generate meaningful and contextually relevant outputs in each area. The strength of our approach lies in its generality, leveraging a variety of data sources such as robotics sequences, gameplay data, large-scale video datasets, and textual information for effective multimodal and multi-task learning. Our approach provides a promising avenue for developing generalist, action-taking, multimodal systems.
format Preprint
id arxiv_https___arxiv_org_abs_2402_05929
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Interactive Agent Foundation Model
Durante, Zane
Sarkar, Bidipta
Gong, Ran
Taori, Rohan
Noda, Yusuke
Tang, Paul
Adeli, Ehsan
Lakshmikanth, Shrinidhi Kowshika
Schulman, Kevin
Milstein, Arnold
Terzopoulos, Demetri
Famoti, Ade
Kuno, Noboru
Llorens, Ashley
Vo, Hoi
Ikeuchi, Katsu
Fei-Fei, Li
Gao, Jianfeng
Wake, Naoki
Huang, Qiuyuan
Artificial Intelligence
Machine Learning
Robotics
The development of artificial intelligence systems is transitioning from creating static, task-specific models to dynamic, agent-based systems capable of performing well in a wide range of applications. We propose an Interactive Agent Foundation Model that uses a novel multi-task agent training paradigm for training AI agents across a wide range of domains, datasets, and tasks. Our training paradigm unifies diverse pre-training strategies, including visual masked auto-encoders, language modeling, and next-action prediction, enabling a versatile and adaptable AI framework. We demonstrate the performance of our framework across three separate domains -- Robotics, Gaming AI, and Healthcare. Our model demonstrates its ability to generate meaningful and contextually relevant outputs in each area. The strength of our approach lies in its generality, leveraging a variety of data sources such as robotics sequences, gameplay data, large-scale video datasets, and textual information for effective multimodal and multi-task learning. Our approach provides a promising avenue for developing generalist, action-taking, multimodal systems.
title An Interactive Agent Foundation Model
topic Artificial Intelligence
Machine Learning
Robotics
url https://arxiv.org/abs/2402.05929