Words in Motion: Extracting Interpretable Control Vectors for Motion Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tas, Omer Sahin, Wagner, Royden
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908367393390592
author Tas, Omer Sahin
Wagner, Royden
author_facet Tas, Omer Sahin
Wagner, Royden
contents Transformer-based models generate hidden states that are difficult to interpret. In this work, we analyze hidden states and modify them at inference, with a focus on motion forecasting. We use linear probing to analyze whether interpretable features are embedded in hidden states. Our experiments reveal high probing accuracy, indicating latent space regularities with functionally important directions. Building on this, we use the directions between hidden states with opposing features to fit control vectors. At inference, we add our control vectors to hidden states and evaluate their impact on predictions. Remarkably, such modifications preserve the feasibility of predictions. We further refine our control vectors using sparse autoencoders (SAEs). This leads to more linear changes in predictions when scaling control vectors. Our approach enables mechanistic interpretation as well as zero-shot generalization to unseen dataset characteristics with negligible computational overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11624
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Words in Motion: Extracting Interpretable Control Vectors for Motion Transformers
Tas, Omer Sahin
Wagner, Royden
Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
Transformer-based models generate hidden states that are difficult to interpret. In this work, we analyze hidden states and modify them at inference, with a focus on motion forecasting. We use linear probing to analyze whether interpretable features are embedded in hidden states. Our experiments reveal high probing accuracy, indicating latent space regularities with functionally important directions. Building on this, we use the directions between hidden states with opposing features to fit control vectors. At inference, we add our control vectors to hidden states and evaluate their impact on predictions. Remarkably, such modifications preserve the feasibility of predictions. We further refine our control vectors using sparse autoencoders (SAEs). This leads to more linear changes in predictions when scaling control vectors. Our approach enables mechanistic interpretation as well as zero-shot generalization to unseen dataset characteristics with negligible computational overhead.
title Words in Motion: Extracting Interpretable Control Vectors for Motion Transformers
topic Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.11624