Structural Inference: Interpreting Small Language Models with Susceptibilities
Fuente:
arXiv
Saved in:
| Main Authors: | Baker, Garrett, Wang, George, Hoogland, Jesse, Murfet, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Embryology of a Language Model
by: Wang, George, et al.
Published: (2025)
by: Wang, George, et al.
Published: (2025)
Towards Spectroscopy: Susceptibility Clusters in Language Models
by: Gordon, Andrew, et al.
Published: (2026)
by: Gordon, Andrew, et al.
Published: (2026)
Dynamics of Transient Structure in In-Context Linear Regression Transformers
by: Carroll, Liam, et al.
Published: (2025)
by: Carroll, Liam, et al.
Published: (2025)
Patterning: The Dual of Interpretability
by: Wang, George, et al.
Published: (2026)
by: Wang, George, et al.
Published: (2026)
Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient
by: Wang, George, et al.
Published: (2024)
by: Wang, George, et al.
Published: (2024)
Compressibility Measures Complexity: Minimum Description Length Meets Singular Learning Theory
by: Urdshals, Einar, et al.
Published: (2025)
by: Urdshals, Einar, et al.
Published: (2025)
Interpreting Reinforcement Learning Agents with Susceptibilities
by: Elliott, Chris, et al.
Published: (2026)
by: Elliott, Chris, et al.
Published: (2026)
Loss Landscape Degeneracy and Stagewise Development in Transformers
by: Hoogland, Jesse, et al.
Published: (2024)
by: Hoogland, Jesse, et al.
Published: (2024)
You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation
by: Lehalleur, Simon Pepin, et al.
Published: (2025)
by: Lehalleur, Simon Pepin, et al.
Published: (2025)
The Loss Kernel: A Geometric Probe for Deep Learning Interpretability
by: Adam, Maxwell, et al.
Published: (2025)
by: Adam, Maxwell, et al.
Published: (2025)
Susceptibilities and Patterning: A Primer on Linear Response in Bayesian Learning
by: Elliott, Chris, et al.
Published: (2026)
by: Elliott, Chris, et al.
Published: (2026)
From Global to Local: A Scalable Benchmark for Local Posterior Sampling
by: Hitchcock, Rohan, et al.
Published: (2025)
by: Hitchcock, Rohan, et al.
Published: (2025)
Modes of Sequence Models and Learning Coefficients
by: Chen, Zhongtian, et al.
Published: (2025)
by: Chen, Zhongtian, et al.
Published: (2025)
Linear Response Estimators for Singular Statistical Models
by: Elliott, Chris, et al.
Published: (2026)
by: Elliott, Chris, et al.
Published: (2026)
Programs as Singularities
by: Murfet, Daniel, et al.
Published: (2025)
by: Murfet, Daniel, et al.
Published: (2025)
Influence Dynamics and Stagewise Data Attribution
by: Lee, Jin Hwa, et al.
Published: (2025)
by: Lee, Jin Hwa, et al.
Published: (2025)
Generalization Boundaries of Fine-Tuned Small Language Models for Graph Structural Inference
by: Podstawski, Michal
Published: (2026)
by: Podstawski, Michal
Published: (2026)
Bayesian Influence Functions for Hessian-Free Data Attribution
by: Kreer, Philipp Alexander, et al.
Published: (2025)
by: Kreer, Philipp Alexander, et al.
Published: (2025)
Open Problems in Mechanistic Interpretability
by: Sharkey, Lee, et al.
Published: (2025)
by: Sharkey, Lee, et al.
Published: (2025)
The Local Learning Coefficient: A Singularity-Aware Complexity Measure
by: Lau, Edmund, et al.
Published: (2023)
by: Lau, Edmund, et al.
Published: (2023)
Stagewise Reinforcement Learning and the Geometry of the Regret Landscape
by: Elliott, Chris, et al.
Published: (2026)
by: Elliott, Chris, et al.
Published: (2026)
IMU-1: Sample-Efficient Pre-training of Small Language Models
by: Grigorev, George
Published: (2026)
by: Grigorev, George
Published: (2026)
Graph Property Inference in Small Language Models: Effects of Representation and Reasoning Strategy
by: Podstawski, Michal
Published: (2026)
by: Podstawski, Michal
Published: (2026)
Unraveling the Potential of Diffusion Models in Small Molecule Generation
by: Zhang, Peining, et al.
Published: (2025)
by: Zhang, Peining, et al.
Published: (2025)
Multitask learning with semiempirical orbital charges enables sample-efficient MLIPs
by: Neporozhnii, Ihor, et al.
Published: (2026)
by: Neporozhnii, Ihor, et al.
Published: (2026)
Extracting Interpretable Task-Specific Circuits from Large Language Models for Faster Inference
by: García-Carrasco, Jorge, et al.
Published: (2024)
by: García-Carrasco, Jorge, et al.
Published: (2024)
Dystruct: Dynamically Structured Diffusion Language Model Decoding via Bayesian Inference
by: Sun, Bian, et al.
Published: (2026)
by: Sun, Bian, et al.
Published: (2026)
Individual Causal Inference with Structural Causal Model
by: Chang, Daniel T.
Published: (2025)
by: Chang, Daniel T.
Published: (2025)
Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language Models
by: Leask, Patrick, et al.
Published: (2025)
by: Leask, Patrick, et al.
Published: (2025)
Interpretable Relational Inference with LLM-Guided Symbolic Dynamics Modeling
by: Liang, Xiaoxiao, et al.
Published: (2026)
by: Liang, Xiaoxiao, et al.
Published: (2026)
Algorithmic Accountability in Small Data: Sample-Size-Induced Bias Within Classification Metrics
by: Briscoe, Jarren, et al.
Published: (2025)
by: Briscoe, Jarren, et al.
Published: (2025)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
by: Kim, Geonhee, et al.
Published: (2024)
by: Kim, Geonhee, et al.
Published: (2024)
TinyGraphEstimator: Adapting Lightweight Language Models for Graph Structure Inference
by: Podstawski, Michal
Published: (2025)
by: Podstawski, Michal
Published: (2025)
Data Selection: A General Principle for Building Small Interpretable Models
by: Ghose, Abhishek
Published: (2022)
by: Ghose, Abhishek
Published: (2022)
TinyLLM: Evaluation and Optimization of Small Language Models for Agentic Tasks on Edge Devices
by: Haque, Mohd Ariful, et al.
Published: (2025)
by: Haque, Mohd Ariful, et al.
Published: (2025)
Large Language Models for Zero-shot Inference of Causal Structures in Biology
by: Newsham, Izzy, et al.
Published: (2025)
by: Newsham, Izzy, et al.
Published: (2025)
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
ABROCA Distributions For Algorithmic Bias Assessment: Considerations Around Interpretation
by: Borchers, Conrad, et al.
Published: (2024)
by: Borchers, Conrad, et al.
Published: (2024)
Route Sparse Autoencoder to Interpret Large Language Models
by: Shi, Wei, et al.
Published: (2025)
by: Shi, Wei, et al.
Published: (2025)
Domain-Adaptive Small Language Models for Structured Tax Code Prediction
by: Nath, Souvik, et al.
Published: (2025)
by: Nath, Souvik, et al.
Published: (2025)
Similar Items
-
Embryology of a Language Model
by: Wang, George, et al.
Published: (2025) -
Towards Spectroscopy: Susceptibility Clusters in Language Models
by: Gordon, Andrew, et al.
Published: (2026) -
Dynamics of Transient Structure in In-Context Linear Regression Transformers
by: Carroll, Liam, et al.
Published: (2025) -
Patterning: The Dual of Interpretability
by: Wang, George, et al.
Published: (2026) -
Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient
by: Wang, George, et al.
Published: (2024)