Salvato in:
Dettagli Bibliografici
Autori principali: Jucys, Karolis, Adamopoulos, George, Hamidi, Mehrab, Milani, Stephanie, Samsami, Mohammad Reza, Zholus, Artem, Joseph, Sonia, Richards, Blake, Rish, Irina, Şimşek, Özgür
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:https://arxiv.org/abs/2407.12161
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913434322337792
author Jucys, Karolis
Adamopoulos, George
Hamidi, Mehrab
Milani, Stephanie
Samsami, Mohammad Reza
Zholus, Artem
Joseph, Sonia
Richards, Blake
Rish, Irina
Şimşek, Özgür
author_facet Jucys, Karolis
Adamopoulos, George
Hamidi, Mehrab
Milani, Stephanie
Samsami, Mohammad Reza
Zholus, Artem
Joseph, Sonia
Richards, Blake
Rish, Irina
Şimşek, Özgür
contents Understanding the mechanisms behind decisions taken by large foundation models in sequential decision making tasks is critical to ensuring that such systems operate transparently and safely. In this work, we perform exploratory analysis on the Video PreTraining (VPT) Minecraft playing agent, one of the largest open-source vision-based agents. We aim to illuminate its reasoning mechanisms by applying various interpretability techniques. First, we analyze the attention mechanism while the agent solves its training task - crafting a diamond pickaxe. The agent pays attention to the last four frames and several key-frames further back in its six-second memory. This is a possible mechanism for maintaining coherence in a task that takes 3-10 minutes, despite the short memory span. Secondly, we perform various interventions, which help us uncover a worrying case of goal misgeneralization: VPT mistakenly identifies a villager wearing brown clothes as a tree trunk when the villager is positioned stationary under green tree leaves, and punches it to death.
format Preprint
id arxiv_https___arxiv_org_abs_2407_12161
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Interpretability in Action: Exploratory Analysis of VPT, a Minecraft Agent
Jucys, Karolis
Adamopoulos, George
Hamidi, Mehrab
Milani, Stephanie
Samsami, Mohammad Reza
Zholus, Artem
Joseph, Sonia
Richards, Blake
Rish, Irina
Şimşek, Özgür
Artificial Intelligence
Understanding the mechanisms behind decisions taken by large foundation models in sequential decision making tasks is critical to ensuring that such systems operate transparently and safely. In this work, we perform exploratory analysis on the Video PreTraining (VPT) Minecraft playing agent, one of the largest open-source vision-based agents. We aim to illuminate its reasoning mechanisms by applying various interpretability techniques. First, we analyze the attention mechanism while the agent solves its training task - crafting a diamond pickaxe. The agent pays attention to the last four frames and several key-frames further back in its six-second memory. This is a possible mechanism for maintaining coherence in a task that takes 3-10 minutes, despite the short memory span. Secondly, we perform various interventions, which help us uncover a worrying case of goal misgeneralization: VPT mistakenly identifies a villager wearing brown clothes as a tree trunk when the villager is positioned stationary under green tree leaves, and punches it to death.
title Interpretability in Action: Exploratory Analysis of VPT, a Minecraft Agent
topic Artificial Intelligence
url https://arxiv.org/abs/2407.12161