Latent Concept Disentanglement in Transformer-based Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hong, Guan Zhe, Vasudeva, Bhavya, Sharan, Vatsal, Rashtchian, Cyrus, Raghavan, Prabhakar, Panigrahy, Rina
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908559939207168
author Hong, Guan Zhe
Vasudeva, Bhavya
Sharan, Vatsal
Rashtchian, Cyrus
Raghavan, Prabhakar
Panigrahy, Rina
author_facet Hong, Guan Zhe
Vasudeva, Bhavya
Sharan, Vatsal
Rashtchian, Cyrus
Raghavan, Prabhakar
Panigrahy, Rina
contents When large language models (LLMs) use in-context learning (ICL) to solve a new task, they must infer latent concepts from demonstration examples. This raises the question of whether and how transformers represent latent structures as part of their computation. Our work experiments with several controlled tasks, studying this question using mechanistic interpretability. First, we show that in transitive reasoning tasks with a latent, discrete concept, the model successfully identifies the latent concept and does step-by-step concept composition. This builds upon prior work that analyzes single-step reasoning. Then, we consider tasks parameterized by a latent numerical concept. We discover low-dimensional subspaces in the model's representation space, where the geometry cleanly reflects the underlying parameterization. Overall, we show that small and large models can indeed disentangle and utilize latent concepts that they learn in-context from a handful of abbreviated demonstrations.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16975
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Latent Concept Disentanglement in Transformer-based Language Models
Hong, Guan Zhe
Vasudeva, Bhavya
Sharan, Vatsal
Rashtchian, Cyrus
Raghavan, Prabhakar
Panigrahy, Rina
Machine Learning
Artificial Intelligence
Computation and Language
When large language models (LLMs) use in-context learning (ICL) to solve a new task, they must infer latent concepts from demonstration examples. This raises the question of whether and how transformers represent latent structures as part of their computation. Our work experiments with several controlled tasks, studying this question using mechanistic interpretability. First, we show that in transitive reasoning tasks with a latent, discrete concept, the model successfully identifies the latent concept and does step-by-step concept composition. This builds upon prior work that analyzes single-step reasoning. Then, we consider tasks parameterized by a latent numerical concept. We discover low-dimensional subspaces in the model's representation space, where the geometry cleanly reflects the underlying parameterization. Overall, we show that small and large models can indeed disentangle and utilize latent concepts that they learn in-context from a handful of abbreviated demonstrations.
title Latent Concept Disentanglement in Transformer-based Language Models
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.16975