A Review of Developmental Interpretability in Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Kendiukhov, Ihor |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Exhaustive Circuit Mapping of a Single-Cell Foundation Model Reveals Massive Redundancy, Heavy-Tailed Hub Architecture, and Layer-Dependent Differentiation Control
par: Kendiukhov, Ihor
Publié: (2026)
par: Kendiukhov, Ihor
Publié: (2026)
Causal Circuit Tracing Reveals Distinct Computational Architectures in Single-Cell Foundation Models: Inhibitory Dominance, Biological Coherence, and Cross-Model Convergence
par: Kendiukhov, Ihor
Publié: (2026)
par: Kendiukhov, Ihor
Publié: (2026)
What Topological and Geometric Structure Do Biological Foundation Models Learn? Evidence from 141 Hypotheses
par: Kendiukhov, Ihor
Publié: (2026)
par: Kendiukhov, Ihor
Publié: (2026)
Discovery of a Hematopoietic Manifold in scGPT Yields a Method for Extracting Performant Algorithms from Biological Foundation Model Internals
par: Kendiukhov, Ihor
Publié: (2026)
par: Kendiukhov, Ihor
Publié: (2026)
Scaling Laws for Masked-Reconstruction Transformers on Single-Cell Transcriptomics
par: Kendiukhov, Ihor
Publié: (2026)
par: Kendiukhov, Ihor
Publié: (2026)
Quantifying Ranking Instability Across Evaluation Protocol Axes in Gene Regulatory Network Benchmarking
par: Kendiukhov, Ihor
Publié: (2026)
par: Kendiukhov, Ihor
Publié: (2026)
Multi-Dimensional Spectral Geometry of Biological Knowledge in Single-Cell Transformer Representations
par: Kendiukhov, Ihor
Publié: (2026)
par: Kendiukhov, Ihor
Publié: (2026)
Sparse autoencoders reveal organized biological knowledge but minimal regulatory logic in single-cell foundation models: a comparative atlas of Geneformer and scGPT
par: Kendiukhov, Ihor
Publié: (2026)
par: Kendiukhov, Ihor
Publié: (2026)
Crafting Large Language Models for Enhanced Interpretability
par: Sun, Chung-En, et autres
Publié: (2024)
par: Sun, Chung-En, et autres
Publié: (2024)
Automatically Interpreting Millions of Features in Large Language Models
par: Paulo, Gonçalo, et autres
Publié: (2024)
par: Paulo, Gonçalo, et autres
Publié: (2024)
Ergodicity Library: A Python Toolkit for Stochastic-Process Simulation, Time-Average Diagnostics, and Agent-Based Experiments
par: Kendiukhov, Ihor
Publié: (2026)
par: Kendiukhov, Ihor
Publié: (2026)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
par: Peng, Runyu, et autres
Publié: (2026)
par: Peng, Runyu, et autres
Publié: (2026)
Efficacy of Large Language Models in Systematic Reviews
par: Shah, Aaditya, et autres
Publié: (2024)
par: Shah, Aaditya, et autres
Publié: (2024)
Rethinking Interpretability in the Era of Large Language Models
par: Singh, Chandan, et autres
Publié: (2024)
par: Singh, Chandan, et autres
Publié: (2024)
A Critical Review of Causal Reasoning Benchmarks for Large Language Models
par: Yang, Linying, et autres
Publié: (2024)
par: Yang, Linying, et autres
Publié: (2024)
Capability-Guided Compression: Toward Interpretability-Aware Budget Allocation for Large Language Models
par: Gupta, Rishaank
Publié: (2026)
par: Gupta, Rishaank
Publié: (2026)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
par: Cho, Hakaze, et autres
Publié: (2025)
par: Cho, Hakaze, et autres
Publié: (2025)
GLiNER-Relex: A Unified Framework for Joint Named Entity Recognition and Relation Extraction
par: Stepanov, Ihor, et autres
Publié: (2026)
par: Stepanov, Ihor, et autres
Publié: (2026)
CrossTrafficLLM: A Human-Centric Framework for Interpretable Traffic Intelligence via Large Language Model
par: Du, Zeming, et autres
Publié: (2025)
par: Du, Zeming, et autres
Publié: (2025)
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content
par: Stepanov, Ihor, et autres
Publié: (2026)
par: Stepanov, Ihor, et autres
Publié: (2026)
Fine-Grained Interpretation of Political Opinions in Large Language Models
par: Hu, Jingyu, et autres
Publié: (2025)
par: Hu, Jingyu, et autres
Publié: (2025)
GLiNER multi-task: Generalist Lightweight Model for Various Information Extraction Tasks
par: Stepanov, Ihor, et autres
Publié: (2024)
par: Stepanov, Ihor, et autres
Publié: (2024)
SelfIE: Self-Interpretation of Large Language Model Embeddings
par: Chen, Haozhe, et autres
Publié: (2024)
par: Chen, Haozhe, et autres
Publié: (2024)
TracrBench: Generating Interpretability Testbeds with Large Language Models
par: Thurnherr, Hannes, et autres
Publié: (2024)
par: Thurnherr, Hannes, et autres
Publié: (2024)
Large Language Models For Text Classification: Case Study And Comprehensive Review
par: Kostina, Arina, et autres
Publié: (2025)
par: Kostina, Arina, et autres
Publié: (2025)
A Review on Scientific Knowledge Extraction using Large Language Models in Biomedical Sciences
par: Garcia, Gabriel Lino, et autres
Publié: (2024)
par: Garcia, Gabriel Lino, et autres
Publié: (2024)
Systematic Evaluation of Single-Cell Foundation Model Interpretability Reveals Attention Captures Co-Expression Rather Than Unique Regulatory Signal
par: Kendiukhov, Ihor
Publié: (2026)
par: Kendiukhov, Ihor
Publié: (2026)
TeleTables: A Benchmark for Large Language Models in Telecom Table Interpretation
par: Ezzakri, Anas, et autres
Publié: (2025)
par: Ezzakri, Anas, et autres
Publié: (2025)
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
par: Shu, Dong, et autres
Publié: (2025)
par: Shu, Dong, et autres
Publié: (2025)
Focus-LIME: Surgical Interpretation of Long-Context Large Language Models via Proxy-Based Neighborhood Selection
par: Liu, Junhao, et autres
Publié: (2026)
par: Liu, Junhao, et autres
Publié: (2026)
The Million-Label NER: Breaking Scale Barriers with GLiNER bi-encoder
par: Stepanov, Ihor, et autres
Publié: (2026)
par: Stepanov, Ihor, et autres
Publié: (2026)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
par: Soo, Samuel, et autres
Publié: (2025)
par: Soo, Samuel, et autres
Publié: (2025)
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
par: Karvonen, Adam, et autres
Publié: (2025)
par: Karvonen, Adam, et autres
Publié: (2025)
How do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective
par: Liu, Qi, et autres
Publié: (2025)
par: Liu, Qi, et autres
Publié: (2025)
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
par: Laptev, Daniil, et autres
Publié: (2025)
par: Laptev, Daniil, et autres
Publié: (2025)
Improving Neuron-level Interpretability with White-box Language Models
par: Bai, Hao, et autres
Publié: (2024)
par: Bai, Hao, et autres
Publié: (2024)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
par: Huang, Jing, et autres
Publié: (2024)
par: Huang, Jing, et autres
Publié: (2024)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
par: Lan, Michael, et autres
Publié: (2023)
par: Lan, Michael, et autres
Publié: (2023)
Interpreting the Repeated Token Phenomenon in Large Language Models
par: Yona, Itay, et autres
Publié: (2025)
par: Yona, Itay, et autres
Publié: (2025)
Towards Incremental Learning in Large Language Models: A Critical Review
par: Jovanovic, Mladjan, et autres
Publié: (2024)
par: Jovanovic, Mladjan, et autres
Publié: (2024)
Documents similaires
-
Exhaustive Circuit Mapping of a Single-Cell Foundation Model Reveals Massive Redundancy, Heavy-Tailed Hub Architecture, and Layer-Dependent Differentiation Control
par: Kendiukhov, Ihor
Publié: (2026) -
Causal Circuit Tracing Reveals Distinct Computational Architectures in Single-Cell Foundation Models: Inhibitory Dominance, Biological Coherence, and Cross-Model Convergence
par: Kendiukhov, Ihor
Publié: (2026) -
What Topological and Geometric Structure Do Biological Foundation Models Learn? Evidence from 141 Hypotheses
par: Kendiukhov, Ihor
Publié: (2026) -
Discovery of a Hematopoietic Manifold in scGPT Yields a Method for Extracting Performant Algorithms from Biological Foundation Model Internals
par: Kendiukhov, Ihor
Publié: (2026) -
Scaling Laws for Masked-Reconstruction Transformers on Single-Cell Transcriptomics
par: Kendiukhov, Ihor
Publié: (2026)