LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss
Fuente:
arXiv
Saved in:
| Main Authors: | Rodriguez, Pau, Klein, Michal, Gualdoni, Eleonora, Maiorca, Valentino, Blaas, Arno, Zappella, Luca, Cuturi, Marco, Suau, Xavier |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HyperTransport: Amortized Conditioning of T2I Generative Models
by: Maiorca, Valentino, et al.
Published: (2026)
by: Maiorca, Valentino, et al.
Published: (2026)
Controlling Language and Diffusion Models by Transporting Activations
by: Rodriguez, Pau, et al.
Published: (2024)
by: Rodriguez, Pau, et al.
Published: (2024)
GenCtrl -- A Formal Controllability Toolkit for Generative Models
by: Cheng, Emily, et al.
Published: (2026)
by: Cheng, Emily, et al.
Published: (2026)
Interpreting CLIP: Insights on the Robustness to ImageNet Distribution Shifts
by: Crabbé, Jonathan, et al.
Published: (2023)
by: Crabbé, Jonathan, et al.
Published: (2023)
Dynamically Scaled Activation Steering
by: Ferrando, Alex, et al.
Published: (2025)
by: Ferrando, Alex, et al.
Published: (2025)
Sparse Autoencoders are Capable LLM Jailbreak Mitigators
by: Assogba, Yannick, et al.
Published: (2026)
by: Assogba, Yannick, et al.
Published: (2026)
ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language Models
by: Danieli, Federico, et al.
Published: (2025)
by: Danieli, Federico, et al.
Published: (2025)
Whispering Experts: Neural Interventions for Toxicity Mitigation in Language Models
by: Suau, Xavier, et al.
Published: (2024)
by: Suau, Xavier, et al.
Published: (2024)
Considerations for Distribution Shift Robustness of Diagnostic Models in Healthcare
by: Blaas, Arno, et al.
Published: (2024)
by: Blaas, Arno, et al.
Published: (2024)
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
by: Bethune, Louis, et al.
Published: (2025)
by: Bethune, Louis, et al.
Published: (2025)
DynaMiCS: Fine-tuning LLMs with Performance Constraints using Dynamic Mixtures
by: Gualdoni, Eleonora, et al.
Published: (2026)
by: Gualdoni, Eleonora, et al.
Published: (2026)
Why do objects have many names? A study on word informativeness in language use and lexical systems
by: Gualdoni, Eleonora, et al.
Published: (2024)
by: Gualdoni, Eleonora, et al.
Published: (2024)
Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results
by: Santilli, Andrea, et al.
Published: (2025)
by: Santilli, Andrea, et al.
Published: (2025)
Uncertainty Quantification for LLM Function-Calling
by: Ye, Zihuiwen, et al.
Published: (2026)
by: Ye, Zihuiwen, et al.
Published: (2026)
Nectar: Neural Estimation of Cached-Token Attention via Regression
by: Monteiro, João, et al.
Published: (2026)
by: Monteiro, João, et al.
Published: (2026)
Multivariate Conformal Prediction using Optimal Transport
by: Klein, Michal, et al.
Published: (2025)
by: Klein, Michal, et al.
Published: (2025)
Contrasting Multiple Representations with the Multi-Marginal Matching Gap
by: Piran, Zoe, et al.
Published: (2024)
by: Piran, Zoe, et al.
Published: (2024)
Latent Space Translation via Inverse Relative Projection
by: Maiorca, Valentino, et al.
Published: (2024)
by: Maiorca, Valentino, et al.
Published: (2024)
From Bricks to Bridges: Product of Invariances to Enhance Latent Space Communication
by: Cannistraci, Irene, et al.
Published: (2023)
by: Cannistraci, Irene, et al.
Published: (2023)
On Fitting Flow Models with Large Sinkhorn Couplings
by: Zhang, Stephen, et al.
Published: (2025)
by: Zhang, Stephen, et al.
Published: (2025)
Amortizing Maximum Inner Product Search with Learned Support Functions
by: Olausson, Theo X., et al.
Published: (2026)
by: Olausson, Theo X., et al.
Published: (2026)
Flow Matching with Semidiscrete Couplings
by: Mousavi-Hosseini, Alireza, et al.
Published: (2025)
by: Mousavi-Hosseini, Alireza, et al.
Published: (2025)
Latent Space Translation via Semantic Alignment
by: Maiorca, Valentino, et al.
Published: (2023)
by: Maiorca, Valentino, et al.
Published: (2023)
Latent Functional Maps: a spectral framework for representation alignment
by: Fumero, Marco, et al.
Published: (2024)
by: Fumero, Marco, et al.
Published: (2024)
ExpertLens: Activation steering features are highly interpretable
by: Fedzechkina, Masha, et al.
Published: (2025)
by: Fedzechkina, Masha, et al.
Published: (2025)
ResiDual Transformer Alignment with Spectral Decomposition
by: Basile, Lorenzo, et al.
Published: (2024)
by: Basile, Lorenzo, et al.
Published: (2024)
Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity
by: Huang, Ningyuan, et al.
Published: (2025)
by: Huang, Ningyuan, et al.
Published: (2025)
What do your logits know? (The answer may surprise you!)
by: Fedzechkina, Masha, et al.
Published: (2026)
by: Fedzechkina, Masha, et al.
Published: (2026)
Mapping representations in Reinforcement Learning via Semantic Alignment for Zero-Shot Stitching
by: Ricciardi, Antonio Pio, et al.
Published: (2025)
by: Ricciardi, Antonio Pio, et al.
Published: (2025)
R3L: Relative Representations for Reinforcement Learning
by: Ricciardi, Antonio Pio, et al.
Published: (2024)
by: Ricciardi, Antonio Pio, et al.
Published: (2024)
Attention to Mamba: A Recipe for Cross-Architecture Distillation
by: Moudgil, Abhinav, et al.
Published: (2026)
by: Moudgil, Abhinav, et al.
Published: (2026)
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
by: Kirchhof, Michael, et al.
Published: (2025)
by: Kirchhof, Michael, et al.
Published: (2025)
Mitigating Content Effects on Reasoning in Language Models through Fine-Grained Activation Steering
by: Valentino, Marco, et al.
Published: (2025)
by: Valentino, Marco, et al.
Published: (2025)
GENOT: Entropic (Gromov) Wasserstein Flow Matching with Applications to Single-Cell Genomics
by: Klein, Dominik, et al.
Published: (2023)
by: Klein, Dominik, et al.
Published: (2023)
Metric Based Few-Shot Graph Classification
by: Crisostomi, Donato, et al.
Published: (2022)
by: Crisostomi, Donato, et al.
Published: (2022)
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
by: Saada, Thiziri Nait, et al.
Published: (2025)
by: Saada, Thiziri Nait, et al.
Published: (2025)
Careful with that Scalpel: Improving Gradient Surgery with an EMA
by: Hsieh, Yu-Guan, et al.
Published: (2024)
by: Hsieh, Yu-Guan, et al.
Published: (2024)
On a Neural Implementation of Brenier's Polar Factorization
by: Vesseron, Nina, et al.
Published: (2024)
by: Vesseron, Nina, et al.
Published: (2024)
DSO: Direct Steering Optimization for Bias Mitigation
by: Paes, Lucas Monteiro, et al.
Published: (2025)
by: Paes, Lucas Monteiro, et al.
Published: (2025)
The Design Space of Tri-Modal Masked Diffusion Models
by: Bethune, Louis, et al.
Published: (2026)
by: Bethune, Louis, et al.
Published: (2026)
Similar Items
-
HyperTransport: Amortized Conditioning of T2I Generative Models
by: Maiorca, Valentino, et al.
Published: (2026) -
Controlling Language and Diffusion Models by Transporting Activations
by: Rodriguez, Pau, et al.
Published: (2024) -
GenCtrl -- A Formal Controllability Toolkit for Generative Models
by: Cheng, Emily, et al.
Published: (2026) -
Interpreting CLIP: Insights on the Robustness to ImageNet Distribution Shifts
by: Crabbé, Jonathan, et al.
Published: (2023) -
Dynamically Scaled Activation Steering
by: Ferrando, Alex, et al.
Published: (2025)