Saved in:
Bibliographic Details
Main Authors: Kolbeinsson, Arinbjorn, O'Brien, Kyle, Huang, Tianjin, Gao, Shanghua, Liu, Shiwei, Schwarz, Jonathan Richard, Vaidya, Anurag, Mahmood, Faisal, Zitnik, Marinka, Chen, Tianlong, Hartvigsen, Thomas
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2407.06483
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913736846999552
author Kolbeinsson, Arinbjorn
O'Brien, Kyle
Huang, Tianjin
Gao, Shanghua
Liu, Shiwei
Schwarz, Jonathan Richard
Vaidya, Anurag
Mahmood, Faisal
Zitnik, Marinka
Chen, Tianlong
Hartvigsen, Thomas
author_facet Kolbeinsson, Arinbjorn
O'Brien, Kyle
Huang, Tianjin
Gao, Shanghua
Liu, Shiwei
Schwarz, Jonathan Richard
Vaidya, Anurag
Mahmood, Faisal
Zitnik, Marinka
Chen, Tianlong
Hartvigsen, Thomas
contents Test-time interventions for language models can enhance factual accuracy, mitigate harmful outputs, and improve model efficiency without costly retraining. But despite a flood of new methods, different types of interventions are largely developing independently. In practice, multiple interventions must be applied sequentially to the same model, yet we lack standardized ways to study how interventions interact. We fill this gap by introducing composable interventions, a framework to study the effects of using multiple interventions on the same language models, featuring new metrics and a unified codebase. Using our framework, we conduct extensive experiments and compose popular methods from three emerging intervention categories -- Knowledge Editing, Model Compression, and Machine Unlearning. Our results from 310 different compositions uncover meaningful interactions: compression hinders editing and unlearning, composing interventions hinges on their order of application, and popular general-purpose metrics are inadequate for assessing composability. Taken together, our findings showcase clear gaps in composability, suggesting a need for new multi-objective interventions. All of our code is public: https://github.com/hartvigsen-group/composable-interventions.
format Preprint
id arxiv_https___arxiv_org_abs_2407_06483
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Composable Interventions for Language Models
Kolbeinsson, Arinbjorn
O'Brien, Kyle
Huang, Tianjin
Gao, Shanghua
Liu, Shiwei
Schwarz, Jonathan Richard
Vaidya, Anurag
Mahmood, Faisal
Zitnik, Marinka
Chen, Tianlong
Hartvigsen, Thomas
Machine Learning
Computation and Language
Test-time interventions for language models can enhance factual accuracy, mitigate harmful outputs, and improve model efficiency without costly retraining. But despite a flood of new methods, different types of interventions are largely developing independently. In practice, multiple interventions must be applied sequentially to the same model, yet we lack standardized ways to study how interventions interact. We fill this gap by introducing composable interventions, a framework to study the effects of using multiple interventions on the same language models, featuring new metrics and a unified codebase. Using our framework, we conduct extensive experiments and compose popular methods from three emerging intervention categories -- Knowledge Editing, Model Compression, and Machine Unlearning. Our results from 310 different compositions uncover meaningful interactions: compression hinders editing and unlearning, composing interventions hinges on their order of application, and popular general-purpose metrics are inadequate for assessing composability. Taken together, our findings showcase clear gaps in composability, suggesting a need for new multi-objective interventions. All of our code is public: https://github.com/hartvigsen-group/composable-interventions.
title Composable Interventions for Language Models
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2407.06483