Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Rajendran, Goutham, Buchholz, Simon, Aragam, Bryon, Schölkopf, Bernhard, Ravikumar, Pradeep
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912148657012736
author Rajendran, Goutham
Buchholz, Simon
Aragam, Bryon
Schölkopf, Bernhard
Ravikumar, Pradeep
author_facet Rajendran, Goutham
Buchholz, Simon
Aragam, Bryon
Schölkopf, Bernhard
Ravikumar, Pradeep
contents To build intelligent machine learning systems, there are two broad approaches. One approach is to build inherently interpretable models, as endeavored by the growing field of causal representation learning. The other approach is to build highly-performant foundation models and then invest efforts into understanding how they work. In this work, we relate these two approaches and study how to learn human-interpretable concepts from data. Weaving together ideas from both fields, we formally define a notion of concepts and show that they can be provably recovered from diverse data. Experiments on synthetic data and large language models show the utility of our unified approach.
format Preprint
id arxiv_https___arxiv_org_abs_2402_09236
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation Models
Rajendran, Goutham
Buchholz, Simon
Aragam, Bryon
Schölkopf, Bernhard
Ravikumar, Pradeep
Machine Learning
Artificial Intelligence
Statistics Theory
To build intelligent machine learning systems, there are two broad approaches. One approach is to build inherently interpretable models, as endeavored by the growing field of causal representation learning. The other approach is to build highly-performant foundation models and then invest efforts into understanding how they work. In this work, we relate these two approaches and study how to learn human-interpretable concepts from data. Weaving together ideas from both fields, we formally define a notion of concepts and show that they can be provably recovered from diverse data. Experiments on synthetic data and large language models show the utility of our unified approach.
title Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation Models
topic Machine Learning
Artificial Intelligence
Statistics Theory
url https://arxiv.org/abs/2402.09236