Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Laptev, Daniil, Balagansky, Nikita, Aksenov, Yaroslav, Gavrilov, Daniil |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
by: Kurochkin, Vadim, et al.
Published: (2025)
by: Kurochkin, Vadim, et al.
Published: (2025)
Teach Old SAEs New Domain Tricks with Boosting
by: Koriagin, Nikita, et al.
Published: (2025)
by: Koriagin, Nikita, et al.
Published: (2025)
You Do Not Fully Utilize Transformer's Representation Capacity
by: Gerasimov, Gleb, et al.
Published: (2025)
by: Gerasimov, Gleb, et al.
Published: (2025)
Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy
by: Balagansky, Nikita, et al.
Published: (2025)
by: Balagansky, Nikita, et al.
Published: (2025)
Learn Your Reference Model for Real Good Alignment
by: Gorbatovski, Alexey, et al.
Published: (2024)
by: Gorbatovski, Alexey, et al.
Published: (2024)
Linear Transformers with Learnable Kernel Functions are Better In-Context Models
by: Aksenov, Yaroslav, et al.
Published: (2024)
by: Aksenov, Yaroslav, et al.
Published: (2024)
Diffusion Language Models Generation Can Be Halted Early
by: Vaina, Sofia Maria Lo Cicero, et al.
Published: (2023)
by: Vaina, Sofia Maria Lo Cicero, et al.
Published: (2023)
Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors
by: Sinii, Viacheslav, et al.
Published: (2025)
by: Sinii, Viacheslav, et al.
Published: (2025)
Mechanistic Permutability: Match Features Across Layers
by: Balagansky, Nikita, et al.
Published: (2024)
by: Balagansky, Nikita, et al.
Published: (2024)
Next Embedding Prediction Makes World Models Stronger
by: Bredis, George, et al.
Published: (2026)
by: Bredis, George, et al.
Published: (2026)
Steering LLM Reasoning Through Bias-Only Adaptation
by: Sinii, Viacheslav, et al.
Published: (2025)
by: Sinii, Viacheslav, et al.
Published: (2025)
Trust-Region Behavior Blending for On-Policy Distillation
by: Plyusov, Daniil, et al.
Published: (2026)
by: Plyusov, Daniil, et al.
Published: (2026)
Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models
by: Li, Zihao, et al.
Published: (2025)
by: Li, Zihao, et al.
Published: (2025)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
by: Soo, Samuel, et al.
Published: (2025)
by: Soo, Samuel, et al.
Published: (2025)
Guided Star-Shaped Masked Diffusion
by: Meshchaninov, Viacheslav, et al.
Published: (2025)
by: Meshchaninov, Viacheslav, et al.
Published: (2025)
Quantization of Large Language Models with an Overdetermined Basis
by: Merkulov, Daniil, et al.
Published: (2024)
by: Merkulov, Daniil, et al.
Published: (2024)
Test Code Generation for Telecom Software Systems using Two-Stage Generative Model
by: Nabeel, Mohamad, et al.
Published: (2024)
by: Nabeel, Mohamad, et al.
Published: (2024)
nach0-pc: Multi-task Language Model with Molecular Point Cloud Encoder
by: Kuznetsov, Maksim, et al.
Published: (2024)
by: Kuznetsov, Maksim, et al.
Published: (2024)
Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World Success
by: Bredis, George, et al.
Published: (2025)
by: Bredis, George, et al.
Published: (2025)
Automatically Interpreting Millions of Features in Large Language Models
by: Paulo, Gonçalo, et al.
Published: (2024)
by: Paulo, Gonçalo, et al.
Published: (2024)
On Teacher Hacking in Language Model Distillation
by: Tiapkin, Daniil, et al.
Published: (2025)
by: Tiapkin, Daniil, et al.
Published: (2025)
DeMeVa at LeWiDi-2025: Modeling Perspectives with In-Context Learning and Label Distribution Learning
by: Ignatev, Daniil, et al.
Published: (2025)
by: Ignatev, Daniil, et al.
Published: (2025)
Crafting Large Language Models for Enhanced Interpretability
by: Sun, Chung-En, et al.
Published: (2024)
by: Sun, Chung-En, et al.
Published: (2024)
Steering Language Models with Weight Arithmetic
by: Fierro, Constanza, et al.
Published: (2025)
by: Fierro, Constanza, et al.
Published: (2025)
Steering Language Models With Activation Engineering
by: Turner, Alexander Matt, et al.
Published: (2023)
by: Turner, Alexander Matt, et al.
Published: (2023)
When an LLM is apprehensive about its answers -- and when its uncertainty is justified
by: Sychev, Petr, et al.
Published: (2025)
by: Sychev, Petr, et al.
Published: (2025)
Complexity-aware fine-tuning
by: Goncharov, Andrey, et al.
Published: (2025)
by: Goncharov, Andrey, et al.
Published: (2025)
Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models
by: Sundar, Anirudh, et al.
Published: (2025)
by: Sundar, Anirudh, et al.
Published: (2025)
Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention
by: Jin, Zehao, et al.
Published: (2026)
by: Jin, Zehao, et al.
Published: (2026)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
by: He, Zirui, et al.
Published: (2025)
by: He, Zirui, et al.
Published: (2025)
Compositional Steering of Large Language Models with Steering Tokens
by: Radevski, Gorjan, et al.
Published: (2026)
by: Radevski, Gorjan, et al.
Published: (2026)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
by: Lan, Michael, et al.
Published: (2023)
by: Lan, Michael, et al.
Published: (2023)
Differentially Private Steering for Large Language Model Alignment
by: Goel, Anmol, et al.
Published: (2025)
by: Goel, Anmol, et al.
Published: (2025)
Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization
by: Cao, Yuanpu, et al.
Published: (2024)
by: Cao, Yuanpu, et al.
Published: (2024)
EEFSUVA: A New Mathematical Olympiad Benchmark
by: Khatibi, Nicole N, et al.
Published: (2025)
by: Khatibi, Nicole N, et al.
Published: (2025)
Spherical Steering: Geometry-Aware Activation Rotation for Language Models
by: You, Zejia, et al.
Published: (2026)
by: You, Zejia, et al.
Published: (2026)
Disentangling the Roles of Representation and Selection in Data Pruning
by: Du, Yupei, et al.
Published: (2025)
by: Du, Yupei, et al.
Published: (2025)
Superscopes: Amplifying Internal Feature Representations for Language Model Interpretation
by: Jacobi, Jonathan, et al.
Published: (2025)
by: Jacobi, Jonathan, et al.
Published: (2025)
Word Embeddings Are Steers for Language Models
by: Han, Chi, et al.
Published: (2023)
by: Han, Chi, et al.
Published: (2023)
OpenAutoNLU: Open Source AutoML Library for NLU
by: Arshinov, Grigory, et al.
Published: (2026)
by: Arshinov, Grigory, et al.
Published: (2026)
Similar Items
-
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
by: Kurochkin, Vadim, et al.
Published: (2025) -
Teach Old SAEs New Domain Tricks with Boosting
by: Koriagin, Nikita, et al.
Published: (2025) -
You Do Not Fully Utilize Transformer's Representation Capacity
by: Gerasimov, Gleb, et al.
Published: (2025) -
Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy
by: Balagansky, Nikita, et al.
Published: (2025) -
Learn Your Reference Model for Real Good Alignment
by: Gorbatovski, Alexey, et al.
Published: (2024)