From Weights to Activations: Is Steering the Next Frontier of Adaptation?
Fuente:
arXiv
Guardado en:
| Autores principales: | Ostermann, Simon, Gurgurov, Daniil, Baeumel, Tanja, Hedderich, Michael A., Lapuschkin, Sebastian, Samek, Wojciech, Schmitt, Vera |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Sparse Subnetwork Enhancement for Underrepresented Languages in Large Language Models
por: Gurgurov, Daniil, et al.
Publicado: (2025)
por: Gurgurov, Daniil, et al.
Publicado: (2025)
Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection
por: Ghussin, Yusser Al, et al.
Publicado: (2026)
por: Ghussin, Yusser Al, et al.
Publicado: (2026)
Modular Arithmetic: Language Models Solve Math Digit by Digit
por: Baeumel, Tanja, et al.
Publicado: (2025)
por: Baeumel, Tanja, et al.
Publicado: (2025)
CLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark
por: Gurgurov, Daniil, et al.
Publicado: (2026)
por: Gurgurov, Daniil, et al.
Publicado: (2026)
Language Arithmetics: Towards Systematic Language Neuron Identification and Manipulation
por: Gurgurov, Daniil, et al.
Publicado: (2025)
por: Gurgurov, Daniil, et al.
Publicado: (2025)
GrEmLIn: A Repository of Green Baseline Embeddings for 87 Low-Resource Languages Injected with Multilingual Graph Knowledge
por: Gurgurov, Daniil, et al.
Publicado: (2024)
por: Gurgurov, Daniil, et al.
Publicado: (2024)
Adapting Multilingual LLMs to Low-Resource Languages with Knowledge Graphs via Adapters
por: Gurgurov, Daniil, et al.
Publicado: (2024)
por: Gurgurov, Daniil, et al.
Publicado: (2024)
Small Models, Big Impact: Efficient Corpus and Graph-Based Adaptation of Small Multilingual Language Models for Low-Resource Languages
por: Gurgurov, Daniil, et al.
Publicado: (2025)
por: Gurgurov, Daniil, et al.
Publicado: (2025)
Disentangling Mathematical Reasoning in LLMs: A Methodological Investigation of Internal Mechanisms
por: Baeumel, Tanja, et al.
Publicado: (2026)
por: Baeumel, Tanja, et al.
Publicado: (2026)
The Lookahead Limitation: Why Multi-Operand Addition is Hard for LLMs
por: Baeumel, Tanja, et al.
Publicado: (2025)
por: Baeumel, Tanja, et al.
Publicado: (2025)
Multilingual Political Views of Large Language Models: Identification and Steering
por: Gurgurov, Daniil, et al.
Publicado: (2025)
por: Gurgurov, Daniil, et al.
Publicado: (2025)
On Multilingual Encoder Language Model Compression for Low-Resource Languages
por: Gurgurov, Daniil, et al.
Publicado: (2025)
por: Gurgurov, Daniil, et al.
Publicado: (2025)
Multilingual Large Language Models and Curse of Multilinguality
por: Gurgurov, Daniil, et al.
Publicado: (2024)
por: Gurgurov, Daniil, et al.
Publicado: (2024)
DFKI-MLT at SemEval-2026 TASK 7: Steering Multilingual Models Towards Cultural Knowledge
por: Ghussin, Yusser Al, et al.
Publicado: (2026)
por: Ghussin, Yusser Al, et al.
Publicado: (2026)
From Attribution to Action: A Human-Centered Application of Activation Steering
por: Labarta, Tobias, et al.
Publicado: (2026)
por: Labarta, Tobias, et al.
Publicado: (2026)
Atlas-Alignment: Making Interpretability Transferable Across Language Models
por: Puri, Bruno, et al.
Publicado: (2025)
por: Puri, Bruno, et al.
Publicado: (2025)
Circuit Insights: Towards Interpretability Beyond Activations
por: Golimblevskaia, Elena, et al.
Publicado: (2025)
por: Golimblevskaia, Elena, et al.
Publicado: (2025)
Judge Circuits
por: Feldhus, Nils, et al.
Publicado: (2026)
por: Feldhus, Nils, et al.
Publicado: (2026)
The Latin Substrate: How Language Models Represent and Mediate Script Choice
por: Gurgurov, Daniil, et al.
Publicado: (2026)
por: Gurgurov, Daniil, et al.
Publicado: (2026)
A Close Look at Decomposition-based XAI-Methods for Transformer Language Models
por: Arras, Leila, et al.
Publicado: (2025)
por: Arras, Leila, et al.
Publicado: (2025)
ReasonXL: Shifting LLM Reasoning Language Without Sacrificing Performance
por: Gurgurov, Daniil, et al.
Publicado: (2026)
por: Gurgurov, Daniil, et al.
Publicado: (2026)
The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation
por: Kahardipraja, Patrick, et al.
Publicado: (2025)
por: Kahardipraja, Patrick, et al.
Publicado: (2025)
Image-to-LaTeX Converter for Mathematical Formulas and Text
por: Gurgurov, Daniil, et al.
Publicado: (2024)
por: Gurgurov, Daniil, et al.
Publicado: (2024)
Post-Hoc Concept Disentanglement: From Correlated to Isolated Concept Representations
por: Erogullari, Eren, et al.
Publicado: (2025)
por: Erogullari, Eren, et al.
Publicado: (2025)
FADE: Why Bad Descriptions Happen to Good Features
por: Puri, Bruno, et al.
Publicado: (2025)
por: Puri, Bruno, et al.
Publicado: (2025)
What Language(s) Does Aya-23 Think In? How Multilinguality Affects Internal Language Representations
por: Trinley, Katharina, et al.
Publicado: (2025)
por: Trinley, Katharina, et al.
Publicado: (2025)
Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations
por: Dreyer, Maximilian, et al.
Publicado: (2023)
por: Dreyer, Maximilian, et al.
Publicado: (2023)
Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data
por: Pahde, Frederik, et al.
Publicado: (2025)
por: Pahde, Frederik, et al.
Publicado: (2025)
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
por: Achtibat, Reduan, et al.
Publicado: (2024)
por: Achtibat, Reduan, et al.
Publicado: (2024)
Cross-Refine: Improving Natural Language Explanation Generation by Learning in Tandem
por: Wang, Qianli, et al.
Publicado: (2024)
por: Wang, Qianli, et al.
Publicado: (2024)
Iterative Inference in a Chess-Playing Neural Network
por: Sandmann, Elias, et al.
Publicado: (2025)
por: Sandmann, Elias, et al.
Publicado: (2025)
Attribution-Guided Pruning for Insight and Control: Circuit Discovery and Targeted Correction in Small-scale LLMs
por: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Publicado: (2025)
por: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Publicado: (2025)
Human-Centered Evaluation of XAI Methods
por: Dawoud, Karam, et al.
Publicado: (2023)
por: Dawoud, Karam, et al.
Publicado: (2023)
FitCF: A Framework for Automatic Feature Importance-guided Counterfactual Example Generation
por: Wang, Qianli, et al.
Publicado: (2025)
por: Wang, Qianli, et al.
Publicado: (2025)
Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples
por: Bouanani, Oussama, et al.
Publicado: (2026)
por: Bouanani, Oussama, et al.
Publicado: (2026)
Steer2Edit: From Activation Steering to Component-Level Editing
por: Sun, Chung-En, et al.
Publicado: (2026)
por: Sun, Chung-En, et al.
Publicado: (2026)
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
por: Dreyer, Maximilian, et al.
Publicado: (2024)
por: Dreyer, Maximilian, et al.
Publicado: (2024)
Reactive Model Correction: Mitigating Harm to Task-Relevant Features via Conditional Bias Suppression
por: Bareeva, Dilyara, et al.
Publicado: (2024)
por: Bareeva, Dilyara, et al.
Publicado: (2024)
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
por: Wang, Qianli, et al.
Publicado: (2026)
por: Wang, Qianli, et al.
Publicado: (2026)
Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
por: Wang, Qianli, et al.
Publicado: (2025)
por: Wang, Qianli, et al.
Publicado: (2025)
Ejemplares similares
-
Sparse Subnetwork Enhancement for Underrepresented Languages in Large Language Models
por: Gurgurov, Daniil, et al.
Publicado: (2025) -
Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection
por: Ghussin, Yusser Al, et al.
Publicado: (2026) -
Modular Arithmetic: Language Models Solve Math Digit by Digit
por: Baeumel, Tanja, et al.
Publicado: (2025) -
CLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark
por: Gurgurov, Daniil, et al.
Publicado: (2026) -
Language Arithmetics: Towards Systematic Language Neuron Identification and Manipulation
por: Gurgurov, Daniil, et al.
Publicado: (2025)