Teach Old SAEs New Domain Tricks with Boosting
Fuente:
arXiv
Saved in:
| Main Authors: | Koriagin, Nikita, Aksenov, Yaroslav, Laptev, Daniil, Gerasimov, Gleb, Balagansky, Nikita, Gavrilov, Daniil |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
by: Laptev, Daniil, et al.
Published: (2025)
by: Laptev, Daniil, et al.
Published: (2025)
Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy
by: Balagansky, Nikita, et al.
Published: (2025)
by: Balagansky, Nikita, et al.
Published: (2025)
You Do Not Fully Utilize Transformer's Representation Capacity
by: Gerasimov, Gleb, et al.
Published: (2025)
by: Gerasimov, Gleb, et al.
Published: (2025)
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
by: Kurochkin, Vadim, et al.
Published: (2025)
by: Kurochkin, Vadim, et al.
Published: (2025)
Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors
by: Sinii, Viacheslav, et al.
Published: (2025)
by: Sinii, Viacheslav, et al.
Published: (2025)
Learn Your Reference Model for Real Good Alignment
by: Gorbatovski, Alexey, et al.
Published: (2024)
by: Gorbatovski, Alexey, et al.
Published: (2024)
Linear Transformers with Learnable Kernel Functions are Better In-Context Models
by: Aksenov, Yaroslav, et al.
Published: (2024)
by: Aksenov, Yaroslav, et al.
Published: (2024)
Next Embedding Prediction Makes World Models Stronger
by: Bredis, George, et al.
Published: (2026)
by: Bredis, George, et al.
Published: (2026)
Diffusion Language Models Generation Can Be Halted Early
by: Vaina, Sofia Maria Lo Cicero, et al.
Published: (2023)
by: Vaina, Sofia Maria Lo Cicero, et al.
Published: (2023)
Trust-Region Behavior Blending for On-Policy Distillation
by: Plyusov, Daniil, et al.
Published: (2026)
by: Plyusov, Daniil, et al.
Published: (2026)
Mechanistic Permutability: Match Features Across Layers
by: Balagansky, Nikita, et al.
Published: (2024)
by: Balagansky, Nikita, et al.
Published: (2024)
Steering LLM Reasoning Through Bias-Only Adaptation
by: Sinii, Viacheslav, et al.
Published: (2025)
by: Sinii, Viacheslav, et al.
Published: (2025)
Resa: Transparent Reasoning Models via SAEs
by: Wang, Shangshang, et al.
Published: (2025)
by: Wang, Shangshang, et al.
Published: (2025)
SAEs Are Good for Steering -- If You Select the Right Features
by: Arad, Dana, et al.
Published: (2025)
by: Arad, Dana, et al.
Published: (2025)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
by: Song, Xiangchen, et al.
Published: (2025)
by: Song, Xiangchen, et al.
Published: (2025)
Learning Shortest Paths with Generative Flow Networks
by: Morozov, Nikita, et al.
Published: (2026)
by: Morozov, Nikita, et al.
Published: (2026)
Enhancing LLM Evaluations: The Garbling Trick
by: Bradley, William F.
Published: (2024)
by: Bradley, William F.
Published: (2024)
SLaNC: Static LayerNorm Calibration
by: Salmani, Mahsa, et al.
Published: (2024)
by: Salmani, Mahsa, et al.
Published: (2024)
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
by: Plyusov, Daniil, et al.
Published: (2026)
by: Plyusov, Daniil, et al.
Published: (2026)
Improving GFlowNets with Monte Carlo Tree Search
by: Morozov, Nikita, et al.
Published: (2024)
by: Morozov, Nikita, et al.
Published: (2024)
A New Perspective on Transformers in Online Reinforcement Learning for Continuous Control
by: Kachaev, Nikita, et al.
Published: (2025)
by: Kachaev, Nikita, et al.
Published: (2025)
Developing Lightweight DNN Models With Limited Data For Real-Time Sign Language Recognition
by: Nikitin, Nikita, et al.
Published: (2025)
by: Nikitin, Nikita, et al.
Published: (2025)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
On Teacher Hacking in Language Model Distillation
by: Tiapkin, Daniil, et al.
Published: (2025)
by: Tiapkin, Daniil, et al.
Published: (2025)
Comparing Pre-trained Human Language Models: Is it Better with Human Context as Groups, Individual Traits, or Both?
by: Soni, Nikita, et al.
Published: (2024)
by: Soni, Nikita, et al.
Published: (2024)
Large Human Language Models: A Need and the Challenges
by: Soni, Nikita, et al.
Published: (2023)
by: Soni, Nikita, et al.
Published: (2023)
Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World Success
by: Bredis, George, et al.
Published: (2025)
by: Bredis, George, et al.
Published: (2025)
Finetune Once: Decoupling General & Domain Learning with Dynamic Boosted Annealing
by: Tang, Yang, et al.
Published: (2025)
by: Tang, Yang, et al.
Published: (2025)
Guided Star-Shaped Masked Diffusion
by: Meshchaninov, Viacheslav, et al.
Published: (2025)
by: Meshchaninov, Viacheslav, et al.
Published: (2025)
Exploring the Hidden Capacity of LLMs for One-Step Text Generation
by: Mezentsev, Gleb, et al.
Published: (2025)
by: Mezentsev, Gleb, et al.
Published: (2025)
Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
by: Kachaev, Nikita, et al.
Published: (2025)
by: Kachaev, Nikita, et al.
Published: (2025)
ReflectivePrompt: Reflective evolution in autoprompting algorithms
by: Zhuravlev, Viktor N., et al.
Published: (2025)
by: Zhuravlev, Viktor N., et al.
Published: (2025)
Automatic Prompt Optimization with Prompt Distillation
by: Dyagin, Ernest A., et al.
Published: (2025)
by: Dyagin, Ernest A., et al.
Published: (2025)
SambaLingo: Teaching Large Language Models New Languages
by: Csaki, Zoltan, et al.
Published: (2024)
by: Csaki, Zoltan, et al.
Published: (2024)
Performance Insights-based AI-driven Football Transfer Fee Prediction
by: Sulimov, Daniil
Published: (2024)
by: Sulimov, Daniil
Published: (2024)
Evaluation of LLMs-based Hidden States as Author Representations for Psychological Human-Centered NLP Tasks
by: Soni, Nikita, et al.
Published: (2025)
by: Soni, Nikita, et al.
Published: (2025)
Facilitating large language model Russian adaptation with Learned Embedding Propagation
by: Tikhomirov, Mikhail, et al.
Published: (2024)
by: Tikhomirov, Mikhail, et al.
Published: (2024)
XShare: Collaborative in-Batch Expert Sharing for Faster MoE Inference
by: Vankov, Daniil, et al.
Published: (2026)
by: Vankov, Daniil, et al.
Published: (2026)
Language models in molecular discovery
by: Janakarajan, Nikita, et al.
Published: (2023)
by: Janakarajan, Nikita, et al.
Published: (2023)
BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning
by: Zhang, Beichen, et al.
Published: (2025)
by: Zhang, Beichen, et al.
Published: (2025)
Similar Items
-
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
by: Laptev, Daniil, et al.
Published: (2025) -
Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy
by: Balagansky, Nikita, et al.
Published: (2025) -
You Do Not Fully Utilize Transformer's Representation Capacity
by: Gerasimov, Gleb, et al.
Published: (2025) -
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
by: Kurochkin, Vadim, et al.
Published: (2025) -
Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors
by: Sinii, Viacheslav, et al.
Published: (2025)