Beyond Tokens in Language Models: Interpreting Activations through Text Genre Chunks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Benito-Rodriguez, Éloïse, Urdshals, Einar, Nasufi, Jasmina, Pochinkov, Nicky
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914165733457920
author Benito-Rodriguez, Éloïse
Urdshals, Einar
Nasufi, Jasmina
Pochinkov, Nicky
author_facet Benito-Rodriguez, Éloïse
Urdshals, Einar
Nasufi, Jasmina
Pochinkov, Nicky
contents Understanding Large Language Models (LLMs) is key to ensure their safe and beneficial deployment. This task is complicated by the difficulty of interpretability of LLM structures, and the inability to have all their outputs human-evaluated. In this paper, we present the first step towards a predictive framework, where the genre of a text used to prompt an LLM, is predicted based on its activations. Using Mistral-7B and two datasets, we show that genre can be extracted with F1-scores of up to 98% and 71% using scikit-learn classifiers. Across both datasets, results consistently outperform the control task, providing a proof of concept that text genres can be inferred from LLMs with shallow learning models.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16540
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Tokens in Language Models: Interpreting Activations through Text Genre Chunks
Benito-Rodriguez, Éloïse
Urdshals, Einar
Nasufi, Jasmina
Pochinkov, Nicky
Computation and Language
Machine Learning
cs.LG
Understanding Large Language Models (LLMs) is key to ensure their safe and beneficial deployment. This task is complicated by the difficulty of interpretability of LLM structures, and the inability to have all their outputs human-evaluated. In this paper, we present the first step towards a predictive framework, where the genre of a text used to prompt an LLM, is predicted based on its activations. Using Mistral-7B and two datasets, we show that genre can be extracted with F1-scores of up to 98% and 71% using scikit-learn classifiers. Across both datasets, results consistently outperform the control task, providing a proof of concept that text genres can be inferred from LLMs with shallow learning models.
title Beyond Tokens in Language Models: Interpreting Activations through Text Genre Chunks
topic Computation and Language
Machine Learning
cs.LG
url https://arxiv.org/abs/2511.16540