Representations as Language: An Information-Theoretic Framework for Interpretability
Fuente:
arXiv
Saved in:
| Main Authors: | Conklin, Henry, Smith, Kenny |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Information Structure in Mappings: An Approach to Learning, Representation, and Generalisation
by: Conklin, Henry
Published: (2025)
by: Conklin, Henry
Published: (2025)
Emergent Hierarchical Structure in Large Language Models: An Information-Theoretic Framework for Multi-Scale Representation
by: Zhang, Yukin, et al.
Published: (2025)
by: Zhang, Yukin, et al.
Published: (2025)
An Information-Theoretic Framework for Robust Large Language Model Editing
by: Chen, Qizhou, et al.
Published: (2025)
by: Chen, Qizhou, et al.
Published: (2025)
Compositional Generalization Across Distributional Shifts with Sparse Tree Operations
by: Soulos, Paul, et al.
Published: (2024)
by: Soulos, Paul, et al.
Published: (2024)
Serendipity by Design: Evaluating the Impact of Cross-domain Mappings on Human and LLM Creativity
by: Liu, Qiawen Ella, et al.
Published: (2026)
by: Liu, Qiawen Ella, et al.
Published: (2026)
Pelican Soup Framework: A Theoretical Framework for Language Model Capabilities
by: Chiang, Ting-Rui, et al.
Published: (2024)
by: Chiang, Ting-Rui, et al.
Published: (2024)
On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference
by: Ren, Siyu, et al.
Published: (2024)
by: Ren, Siyu, et al.
Published: (2024)
Multi-Scale Manifold Alignment for Interpreting Large Language Models: A Unified Information-Geometric Framework
by: Zhang, Yukun, et al.
Published: (2025)
by: Zhang, Yukun, et al.
Published: (2025)
Information Flow Routes: Automatically Interpreting Language Models at Scale
by: Ferrando, Javier, et al.
Published: (2024)
by: Ferrando, Javier, et al.
Published: (2024)
Comparing Human and Large Language Model Interpretation of Implicit Information
by: De Santis, Antonio, et al.
Published: (2026)
by: De Santis, Antonio, et al.
Published: (2026)
On Linear Representations and Pretraining Data Frequency in Language Models
by: Merullo, Jack, et al.
Published: (2025)
by: Merullo, Jack, et al.
Published: (2025)
Superscopes: Amplifying Internal Feature Representations for Language Model Interpretation
by: Jacobi, Jonathan, et al.
Published: (2025)
by: Jacobi, Jonathan, et al.
Published: (2025)
LLMD: A Large Language Model for Interpreting Longitudinal Medical Records
by: Porter, Robert, et al.
Published: (2024)
by: Porter, Robert, et al.
Published: (2024)
How Many Features Can a Language Model Store Under the Linear Representation Hypothesis?
by: Garg, Nikhil, et al.
Published: (2026)
by: Garg, Nikhil, et al.
Published: (2026)
On Theoretical Interpretations of Concept-Based In-Context Learning
by: Tang, Huaze, et al.
Published: (2025)
by: Tang, Huaze, et al.
Published: (2025)
Cross-Lingual Transfer and Parameter-Efficient Adaptation in the Turkic Language Family: A Theoretical Framework for Low-Resource Language Models
by: Ibrahimzade, O., et al.
Published: (2026)
by: Ibrahimzade, O., et al.
Published: (2026)
CogniDual Framework: Self-Training Large Language Models within a Dual-System Theoretical Framework for Improving Cognitive Tasks
by: Deng, Yongxin, et al.
Published: (2024)
by: Deng, Yongxin, et al.
Published: (2024)
A Monosemantic Attribution Framework for Stable Interpretability in Clinical Neuroscience Transformer-Based Language Models
by: Mamalakis, Michail, et al.
Published: (2026)
by: Mamalakis, Michail, et al.
Published: (2026)
Information-Theoretic Distillation for Reference-less Summarization
by: Jung, Jaehun, et al.
Published: (2024)
by: Jung, Jaehun, et al.
Published: (2024)
Building Models of Neurological Language
by: Watkins, Henry
Published: (2025)
by: Watkins, Henry
Published: (2025)
CrashSage: A Large Language Model-Centered Framework for Contextual and Interpretable Traffic Crash Analysis
by: Zhen, Hao, et al.
Published: (2025)
by: Zhen, Hao, et al.
Published: (2025)
Probing Ethical Framework Representations in Large Language Models: Structure, Entanglement, and Methodological Challenges
by: Xu, Weilun, et al.
Published: (2026)
by: Xu, Weilun, et al.
Published: (2026)
Evaluating the relationship between regularity and learnability in recursive numeral systems using Reinforcement Learning
by: Silvi, Andrea, et al.
Published: (2026)
by: Silvi, Andrea, et al.
Published: (2026)
Unveiling Language Routing Isolation in Multilingual MoE Models for Interpretable Subnetwork Adaptation
by: Zheng, Kening, et al.
Published: (2026)
by: Zheng, Kening, et al.
Published: (2026)
HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
by: Proskurina, Irina, et al.
Published: (2025)
by: Proskurina, Irina, et al.
Published: (2025)
Information Representation Fairness in Long-Document Embeddings: The Peculiar Interaction of Positional and Language Bias
by: Schuhmacher, Elias, et al.
Published: (2026)
by: Schuhmacher, Elias, et al.
Published: (2026)
Theoretical Foundations and Mitigation of Hallucination in Large Language Models
by: Gumaan, Esmail
Published: (2025)
by: Gumaan, Esmail
Published: (2025)
Interpretability Framework for LLMs in Undergraduate Calculus
by: Dakshit, Sagnik, et al.
Published: (2025)
by: Dakshit, Sagnik, et al.
Published: (2025)
Interpretability of Language Models via Task Spaces
by: Weber, Lucas, et al.
Published: (2024)
by: Weber, Lucas, et al.
Published: (2024)
Activation Scaling for Steering and Interpreting Language Models
by: Stoehr, Niklas, et al.
Published: (2024)
by: Stoehr, Niklas, et al.
Published: (2024)
Semantic Substrate Theory: An Operator-Theoretic Framework for Geometric Semantic Drift
by: Russell, Stephen
Published: (2026)
by: Russell, Stephen
Published: (2026)
Unified Lexical Representation for Interpretable Visual-Language Alignment
by: Li, Yifan, et al.
Published: (2024)
by: Li, Yifan, et al.
Published: (2024)
ARCANE: A Multi-Agent Framework for Interpretable and Configurable Alignment
by: Masters, Charlie, et al.
Published: (2025)
by: Masters, Charlie, et al.
Published: (2025)
Interpreting Public Sentiment in Diplomacy Events: A Counterfactual Analysis Framework Using Large Language Models
by: Ouyang, Leyi
Published: (2025)
by: Ouyang, Leyi
Published: (2025)
Using Large Language Models for the Interpretation of Building Regulations
by: Fuchs, Stefan, et al.
Published: (2024)
by: Fuchs, Stefan, et al.
Published: (2024)
Cognitive BASIC: An In-Model Interpreted Reasoning Language for LLMs
by: Kramer, Oliver
Published: (2025)
by: Kramer, Oliver
Published: (2025)
Mechanistic Interpretability of Emotion Inference in Large Language Models
by: Tak, Ala N., et al.
Published: (2025)
by: Tak, Ala N., et al.
Published: (2025)
Interpretable Emergent Language Using Inter-Agent Transformers
by: Bhardwaj, Mannan
Published: (2025)
by: Bhardwaj, Mannan
Published: (2025)
Deep Natural Language Feature Learning for Interpretable Prediction
by: Urrutia, Felipe, et al.
Published: (2023)
by: Urrutia, Felipe, et al.
Published: (2023)
Can LLMs Narrate Tabular Data? An Evaluation Framework for Natural Language Representations of Text-to-SQL System Outputs
by: Singh, Jyotika, et al.
Published: (2025)
by: Singh, Jyotika, et al.
Published: (2025)
Similar Items
-
Information Structure in Mappings: An Approach to Learning, Representation, and Generalisation
by: Conklin, Henry
Published: (2025) -
Emergent Hierarchical Structure in Large Language Models: An Information-Theoretic Framework for Multi-Scale Representation
by: Zhang, Yukin, et al.
Published: (2025) -
An Information-Theoretic Framework for Robust Large Language Model Editing
by: Chen, Qizhou, et al.
Published: (2025) -
Compositional Generalization Across Distributional Shifts with Sparse Tree Operations
by: Soulos, Paul, et al.
Published: (2024) -
Serendipity by Design: Evaluating the Impact of Cross-domain Mappings on Human and LLM Creativity
by: Liu, Qiawen Ella, et al.
Published: (2026)