Guardado en:
Detalles Bibliográficos
Autores principales: Opiełka, Gustaw, Rosenbusch, Hannes, Stevenson, Claire E.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2503.03666
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917946543046656
author Opiełka, Gustaw
Rosenbusch, Hannes
Stevenson, Claire E.
author_facet Opiełka, Gustaw
Rosenbusch, Hannes
Stevenson, Claire E.
contents Analogical reasoning relies on conceptual abstractions, but it is unclear whether Large Language Models (LLMs) harbor such internal representations. We explore distilled representations from LLM activations and find that function vectors (FVs; Todd et al., 2024) - compact representations for in-context learning (ICL) tasks - are not invariant to simple input changes (e.g., open-ended vs. multiple-choice), suggesting they capture more than pure concepts. Using representational similarity analysis (RSA), we localize a small set of attention heads that encode invariant concept vectors (CVs) for verbal concepts like "antonym". These CVs function as feature detectors that operate independently of the final output - meaning that a model may form a correct internal representation yet still produce an incorrect output. Furthermore, CVs can be used to causally guide model behaviour. However, for more abstract concepts like "previous" and "next", we do not observe invariant linear representations, a finding we link to generalizability issues LLMs display within these domains.
format Preprint
id arxiv_https___arxiv_org_abs_2503_03666
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Analogical Reasoning Inside Large Language Models: Concept Vectors and the Limits of Abstraction
Opiełka, Gustaw
Rosenbusch, Hannes
Stevenson, Claire E.
Computation and Language
Machine Learning
Analogical reasoning relies on conceptual abstractions, but it is unclear whether Large Language Models (LLMs) harbor such internal representations. We explore distilled representations from LLM activations and find that function vectors (FVs; Todd et al., 2024) - compact representations for in-context learning (ICL) tasks - are not invariant to simple input changes (e.g., open-ended vs. multiple-choice), suggesting they capture more than pure concepts. Using representational similarity analysis (RSA), we localize a small set of attention heads that encode invariant concept vectors (CVs) for verbal concepts like "antonym". These CVs function as feature detectors that operate independently of the final output - meaning that a model may form a correct internal representation yet still produce an incorrect output. Furthermore, CVs can be used to causally guide model behaviour. However, for more abstract concepts like "previous" and "next", we do not observe invariant linear representations, a finding we link to generalizability issues LLMs display within these domains.
title Analogical Reasoning Inside Large Language Models: Concept Vectors and the Limits of Abstraction
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2503.03666