Emergence of Linear Truth Encodings in Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ravfogel, Shauli, Yehudai, Gilad, Linzen, Tal, Bruna, Joan, Bietti, Alberto |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Geometric Factual Recall in Transformers
par: Ravfogel, Shauli, et autres
Publié: (2026)
par: Ravfogel, Shauli, et autres
Publié: (2026)
RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context
par: Petty, Jackson, et autres
Publié: (2025)
par: Petty, Jackson, et autres
Publié: (2025)
Can LLMs Introspect? A Reality Check
par: Singh, Shashwat, et autres
Publié: (2026)
par: Singh, Shashwat, et autres
Publié: (2026)
Compositional Reasoning with Transformers, RNNs, and Chain of Thought
par: Yehudai, Gilad, et autres
Publié: (2025)
par: Yehudai, Gilad, et autres
Publié: (2025)
Linear Adversarial Concept Erasure
par: Ravfogel, Shauli, et autres
Publié: (2022)
par: Ravfogel, Shauli, et autres
Publié: (2022)
Do Language Models' Words Refer?
par: Mandelkern, Matthew, et autres
Publié: (2023)
par: Mandelkern, Matthew, et autres
Publié: (2023)
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
par: Ben-Zaken, Elad, et autres
Publié: (2021)
par: Ben-Zaken, Elad, et autres
Publié: (2021)
Gumbel Counterfactual Generation From Language Models
par: Ravfogel, Shauli, et autres
Publié: (2024)
par: Ravfogel, Shauli, et autres
Publié: (2024)
Log-linear Guardedness and its Implications
par: Ravfogel, Shauli, et autres
Publié: (2022)
par: Ravfogel, Shauli, et autres
Publié: (2022)
Entailment Semantics Can Be Extracted from an Ideal Language Model
par: Merrill, William, et autres
Publié: (2022)
par: Merrill, William, et autres
Publié: (2022)
From Directions to Regions: Decomposing Activations in Language Models via Local Geometry
par: Shafran, Or, et autres
Publié: (2026)
par: Shafran, Or, et autres
Publié: (2026)
Distributional Associations vs In-Context Reasoning: A Study of Feed-forward and Attention Layers
par: Chen, Lei, et autres
Publié: (2024)
par: Chen, Lei, et autres
Publié: (2024)
The Medium Is Not the Message: Deconfounding Document Embeddings via Linear Concept Erasure
par: Fan, Yu, et autres
Publié: (2025)
par: Fan, Yu, et autres
Publié: (2025)
SPAWNing Structural Priming Predictions from a Cognitively Motivated Parser
par: Prasad, Grusha, et autres
Publié: (2024)
par: Prasad, Grusha, et autres
Publié: (2024)
A Practical Method for Generating String Counterfactuals
par: Avitan, Matan, et autres
Publié: (2024)
par: Avitan, Matan, et autres
Publié: (2024)
State over Tokens: Characterizing the Role of Reasoning Tokens
par: Levy, Mosh, et autres
Publié: (2025)
par: Levy, Mosh, et autres
Publié: (2025)
Kernelized Concept Erasure
par: Ravfogel, Shauli, et autres
Publié: (2022)
par: Ravfogel, Shauli, et autres
Publié: (2022)
The Role of Language Imbalance in Cross-lingual Generalisation: Insights from Cloned Language Experiments
par: Schäfer, Anton, et autres
Publié: (2024)
par: Schäfer, Anton, et autres
Publié: (2024)
How Does Code Pretraining Affect Language Model Task Performance?
par: Petty, Jackson, et autres
Publié: (2024)
par: Petty, Jackson, et autres
Publié: (2024)
Beyond Single Embeddings: Capturing Diverse Targets with Multi-Query Retrieval
par: Chen, Hung-Ting, et autres
Publié: (2025)
par: Chen, Hung-Ting, et autres
Publié: (2025)
Diversity Over Quantity: A Lesson From Few Shot Relation Classification
par: Cohen, Amir DN, et autres
Publié: (2024)
par: Cohen, Amir DN, et autres
Publié: (2024)
IQ Test for LLMs: An Evaluation Framework for Uncovering Core Skills in LLMs
par: Maimon, Aviya, et autres
Publié: (2025)
par: Maimon, Aviya, et autres
Publié: (2025)
Intrinsic Test of Unlearning Using Parametric Knowledge Traces
par: Hong, Yihuai, et autres
Publié: (2024)
par: Hong, Yihuai, et autres
Publié: (2024)
Deconstructing sentence disambiguation by joint latent modeling of reading paradigms: LLM surprisal is not enough
par: Paape, Dario, et autres
Publié: (2026)
par: Paape, Dario, et autres
Publié: (2026)
Manipulating language models' training data to study syntactic constraint learning: the case of English passivization
par: Leong, Cara Su-Yi, et autres
Publié: (2024)
par: Leong, Cara Su-Yi, et autres
Publié: (2024)
Why are language models less surprised than humans? Testing the Parse Multiplicity Mismatch Hypothesis
par: Timkey, William, et autres
Publié: (2026)
par: Timkey, William, et autres
Publié: (2026)
Language Models Struggle to Use Representations Learned In-Context
par: Lepori, Michael A., et autres
Publié: (2026)
par: Lepori, Michael A., et autres
Publié: (2026)
Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map Alignment
par: Rassin, Royi, et autres
Publié: (2023)
par: Rassin, Royi, et autres
Publié: (2023)
Evaluating In-Context Translation with Synchronous Context-Free Grammar Transduction
par: Petty, Jackson, et autres
Publié: (2026)
par: Petty, Jackson, et autres
Publié: (2026)
Representation Surgery: Theory and Practice of Affine Steering
par: Singh, Shashwat, et autres
Publié: (2024)
par: Singh, Shashwat, et autres
Publié: (2024)
Description-Based Text Similarity
par: Ravfogel, Shauli, et autres
Publié: (2023)
par: Ravfogel, Shauli, et autres
Publié: (2023)
The Impact of Depth on Compositional Generalization in Transformer Language Models
par: Petty, Jackson, et autres
Publié: (2023)
par: Petty, Jackson, et autres
Publié: (2023)
What Goes Into a LM Acceptability Judgment? Rethinking the Impact of Frequency and Length
par: Tjuatja, Lindia, et autres
Publié: (2024)
par: Tjuatja, Lindia, et autres
Publié: (2024)
In-context Learning Generalizes, But Not Always Robustly: The Case of Syntax
par: Mueller, Aaron, et autres
Publié: (2023)
par: Mueller, Aaron, et autres
Publié: (2023)
Multilingual Prompting for Improving LLM Generation Diversity
par: Wang, Qihan, et autres
Publié: (2025)
par: Wang, Qihan, et autres
Publié: (2025)
LEACE: Perfect linear concept erasure in closed form
par: Belrose, Nora, et autres
Publié: (2023)
par: Belrose, Nora, et autres
Publié: (2023)
Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models
par: Qiu, Linlu, et autres
Publié: (2025)
par: Qiu, Linlu, et autres
Publié: (2025)
The Truthfulness Spectrum Hypothesis
par: Ying, Zhuofan Josh, et autres
Publié: (2026)
par: Ying, Zhuofan Josh, et autres
Publié: (2026)
On the Benefits of Rank in Attention Layers
par: Amsel, Noah, et autres
Publié: (2024)
par: Amsel, Noah, et autres
Publié: (2024)
Can You Learn Semantics Through Next-Word Prediction? The Case of Entailment
par: Merrill, William, et autres
Publié: (2024)
par: Merrill, William, et autres
Publié: (2024)
Documents similaires
-
Geometric Factual Recall in Transformers
par: Ravfogel, Shauli, et autres
Publié: (2026) -
RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context
par: Petty, Jackson, et autres
Publié: (2025) -
Can LLMs Introspect? A Reality Check
par: Singh, Shashwat, et autres
Publié: (2026) -
Compositional Reasoning with Transformers, RNNs, and Chain of Thought
par: Yehudai, Gilad, et autres
Publié: (2025) -
Linear Adversarial Concept Erasure
par: Ravfogel, Shauli, et autres
Publié: (2022)