The Information Geometry of Softmax: Probing and Steering
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Park, Kiho, Nief, Todd, Choe, Yo Joong, Veitch, Victor |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The Linear Representation Hypothesis and the Geometry of Large Language Models
par: Park, Kiho, et autres
Publié: (2023)
par: Park, Kiho, et autres
Publié: (2023)
The Geometry of Categorical and Hierarchical Concepts in Large Language Models
par: Park, Kiho, et autres
Publié: (2024)
par: Park, Kiho, et autres
Publié: (2024)
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
par: Muchane, Mark, et autres
Publié: (2025)
par: Muchane, Mark, et autres
Publié: (2025)
RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals
par: Reber, David, et autres
Publié: (2024)
par: Reber, David, et autres
Publié: (2024)
Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry
par: Sim, Woo Seob, et autres
Publié: (2026)
par: Sim, Woo Seob, et autres
Publié: (2026)
Scalable-Softmax Is Superior for Attention
par: Nakanishi, Ken M.
Publié: (2025)
par: Nakanishi, Ken M.
Publié: (2025)
To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models
par: Gonsior, Julius, et autres
Publié: (2022)
par: Gonsior, Julius, et autres
Publié: (2022)
Forgetting Transformer: Softmax Attention with a Forget Gate
par: Lin, Zhixuan, et autres
Publié: (2025)
par: Lin, Zhixuan, et autres
Publié: (2025)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
par: Siu, Vincent, et autres
Publié: (2025)
par: Siu, Vincent, et autres
Publié: (2025)
Steer LLM Latents for Hallucination Detection
par: Park, Seongheon, et autres
Publié: (2025)
par: Park, Seongheon, et autres
Publié: (2025)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
par: Collins, Liam, et autres
Publié: (2024)
par: Collins, Liam, et autres
Publié: (2024)
Extracting Unlearned Information from LLMs with Activation Steering
par: Seyitoğlu, Atakan, et autres
Publié: (2024)
par: Seyitoğlu, Atakan, et autres
Publié: (2024)
HyperSteer: Activation Steering at Scale with Hypernetworks
par: Sun, Jiuding, et autres
Publié: (2025)
par: Sun, Jiuding, et autres
Publié: (2025)
Steer Like the LLM: Activation Steering that Mimics Prompting
par: Heyman, Geert, et autres
Publié: (2026)
par: Heyman, Geert, et autres
Publié: (2026)
Compositional Steering of Large Language Models with Steering Tokens
par: Radevski, Gorjan, et autres
Publié: (2026)
par: Radevski, Gorjan, et autres
Publié: (2026)
Universal Conceptual Structure in Neural Translation: Probing NLLB-200's Multilingual Geometry
par: Mathewson, Kyle Elliott
Publié: (2026)
par: Mathewson, Kyle Elliott
Publié: (2026)
Salient Information Prompting to Steer Content in Prompt-based Abstractive Summarization
par: Xu, Lei, et autres
Publié: (2024)
par: Xu, Lei, et autres
Publié: (2024)
Probing Geometry of Next Token Prediction Using Cumulant Expansion of the Softmax Entropy
par: Viswanathan, Karthik, et autres
Publié: (2025)
par: Viswanathan, Karthik, et autres
Publié: (2025)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
par: Cheng, Stephen, et autres
Publié: (2026)
par: Cheng, Stephen, et autres
Publié: (2026)
Steer2Adapt: Dynamically Composing Steering Vectors Elicits Efficient Adaptation of LLMs
par: Han, Pengrui, et autres
Publié: (2026)
par: Han, Pengrui, et autres
Publié: (2026)
COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics
par: Sharma, Kartik, et autres
Publié: (2026)
par: Sharma, Kartik, et autres
Publié: (2026)
FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models
par: Weng, Zixuan, et autres
Publié: (2026)
par: Weng, Zixuan, et autres
Publié: (2026)
Steering LLMs for Formal Theorem Proving
par: Kirtania, Shashank, et autres
Publié: (2025)
par: Kirtania, Shashank, et autres
Publié: (2025)
Programming Refusal with Conditional Activation Steering
par: Lee, Bruce W., et autres
Publié: (2024)
par: Lee, Bruce W., et autres
Publié: (2024)
Word Embeddings Are Steers for Language Models
par: Han, Chi, et autres
Publié: (2023)
par: Han, Chi, et autres
Publié: (2023)
SAKE: Steering Activations for Knowledge Editing
par: Scialanga, Marco, et autres
Publié: (2025)
par: Scialanga, Marco, et autres
Publié: (2025)
Integrating Locality-Aware Attention with Transformers for General Geometry PDEs
par: Koh, Minsu, et autres
Publié: (2025)
par: Koh, Minsu, et autres
Publié: (2025)
Subliminal Learning is a LoRA Artifact
par: Nief, Todd, et autres
Publié: (2026)
par: Nief, Todd, et autres
Publié: (2026)
Understanding and Mitigating Dataset Corruption in LLM Steering
par: Anderson, Cullen, et autres
Publié: (2026)
par: Anderson, Cullen, et autres
Publié: (2026)
Endogenous Resistance to Activation Steering in Language Models
par: McKenzie, Alex, et autres
Publié: (2026)
par: McKenzie, Alex, et autres
Publié: (2026)
Focus On This, Not That! Steering LLMs with Adaptive Feature Specification
par: Lamb, Tom A., et autres
Publié: (2024)
par: Lamb, Tom A., et autres
Publié: (2024)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
par: Chalnev, Sviatoslav, et autres
Publié: (2024)
par: Chalnev, Sviatoslav, et autres
Publié: (2024)
Steering Large Language Models for Machine Translation Personalization
par: Scalena, Daniel, et autres
Publié: (2025)
par: Scalena, Daniel, et autres
Publié: (2025)
Brain-Grounded Axes for Reading and Steering LLM States
par: Andric, Sandro
Publié: (2025)
par: Andric, Sandro
Publié: (2025)
SAEs Are Good for Steering -- If You Select the Right Features
par: Arad, Dana, et autres
Publié: (2025)
par: Arad, Dana, et autres
Publié: (2025)
HelpSteer2-Preference: Complementing Ratings with Preferences
par: Wang, Zhilin, et autres
Publié: (2024)
par: Wang, Zhilin, et autres
Publié: (2024)
Steering Llama 2 via Contrastive Activation Addition
par: Panickssery, Nina, et autres
Publié: (2023)
par: Panickssery, Nina, et autres
Publié: (2023)
Steering Language Models Before They Speak: Logit-Level Interventions
par: An, Hyeseon, et autres
Publié: (2026)
par: An, Hyeseon, et autres
Publié: (2026)
DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment
par: Wedgwood, James, et autres
Publié: (2026)
par: Wedgwood, James, et autres
Publié: (2026)
Improving Instruction-Following in Language Models through Activation Steering
par: Stolfo, Alessandro, et autres
Publié: (2024)
par: Stolfo, Alessandro, et autres
Publié: (2024)
Documents similaires
-
The Linear Representation Hypothesis and the Geometry of Large Language Models
par: Park, Kiho, et autres
Publié: (2023) -
The Geometry of Categorical and Hierarchical Concepts in Large Language Models
par: Park, Kiho, et autres
Publié: (2024) -
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
par: Muchane, Mark, et autres
Publié: (2025) -
RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals
par: Reber, David, et autres
Publié: (2024) -
Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry
par: Sim, Woo Seob, et autres
Publié: (2026)