Saved in:
Bibliographic Details
Main Authors: Costarelli, Anthony, Allen, Mat, Field, Severin
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2410.02472
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913573487247360
author Costarelli, Anthony
Allen, Mat
Field, Severin
author_facet Costarelli, Anthony
Allen, Mat
Field, Severin
contents As Large Language Models (LLMs) become increasingly integrated into our daily lives, the potential harms from deceptive behavior underlie the need for faithfully interpreting their decision-making. While traditional probing methods have shown some effectiveness, they remain best for narrowly scoped tasks while more comprehensive explanations are still necessary. To this end, we investigate meta-models-an architecture using a "meta-model" that takes activations from an "input-model" and answers natural language questions about the input-model's behaviors. We evaluate the meta-model's ability to generalize by training them on selected task types and assessing their out-of-distribution performance in deceptive scenarios. Our findings show that meta-models generalize well to out-of-distribution tasks and point towards opportunities for future research in this area. Our code is available at https://github.com/acostarelli/meta-models-public .
format Preprint
id arxiv_https___arxiv_org_abs_2410_02472
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Meta-Models: An Architecture for Decoding LLM Behaviors Through Interpreted Embeddings and Natural Language
Costarelli, Anthony
Allen, Mat
Field, Severin
Machine Learning
Artificial Intelligence
As Large Language Models (LLMs) become increasingly integrated into our daily lives, the potential harms from deceptive behavior underlie the need for faithfully interpreting their decision-making. While traditional probing methods have shown some effectiveness, they remain best for narrowly scoped tasks while more comprehensive explanations are still necessary. To this end, we investigate meta-models-an architecture using a "meta-model" that takes activations from an "input-model" and answers natural language questions about the input-model's behaviors. We evaluate the meta-model's ability to generalize by training them on selected task types and assessing their out-of-distribution performance in deceptive scenarios. Our findings show that meta-models generalize well to out-of-distribution tasks and point towards opportunities for future research in this area. Our code is available at https://github.com/acostarelli/meta-models-public .
title Meta-Models: An Architecture for Decoding LLM Behaviors Through Interpreted Embeddings and Natural Language
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.02472