Spherical Steering: Geometry-Aware Activation Rotation for Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | You, Zejia, Deng, Chunyuan, Chen, Hanjie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Language Models are Symbolic Learners in Arithmetic
von: Deng, Chunyuan, et al.
Veröffentlicht: (2024)
von: Deng, Chunyuan, et al.
Veröffentlicht: (2024)
Steering Information Utility in Key-Value Memory for Language Model Post-Training
von: Deng, Chunyuan, et al.
Veröffentlicht: (2025)
von: Deng, Chunyuan, et al.
Veröffentlicht: (2025)
Steering Language Models With Activation Engineering
von: Turner, Alexander Matt, et al.
Veröffentlicht: (2023)
von: Turner, Alexander Matt, et al.
Veröffentlicht: (2023)
Learning Distribution-Wise Control in Representation Space for Language Models
von: Deng, Chunyuan, et al.
Veröffentlicht: (2025)
von: Deng, Chunyuan, et al.
Veröffentlicht: (2025)
FishBack: Pullback Fisher Geometry for Optimal Activation Steering in Transformers
von: Wang, Sihan, et al.
Veröffentlicht: (2026)
von: Wang, Sihan, et al.
Veröffentlicht: (2026)
The Generalization Ridge: Information Flow in Natural Language Generation
von: Chang, Ruidi, et al.
Veröffentlicht: (2025)
von: Chang, Ruidi, et al.
Veröffentlicht: (2025)
Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention
von: Jin, Zehao, et al.
Veröffentlicht: (2026)
von: Jin, Zehao, et al.
Veröffentlicht: (2026)
Endogenous Resistance to Activation Steering in Language Models
von: McKenzie, Alex, et al.
Veröffentlicht: (2026)
von: McKenzie, Alex, et al.
Veröffentlicht: (2026)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
von: Soo, Samuel, et al.
Veröffentlicht: (2025)
von: Soo, Samuel, et al.
Veröffentlicht: (2025)
Improving Instruction-Following in Language Models through Activation Steering
von: Stolfo, Alessandro, et al.
Veröffentlicht: (2024)
von: Stolfo, Alessandro, et al.
Veröffentlicht: (2024)
ByteFlow: Language Modeling through Adaptive Byte Compression without a Tokenizer
von: Deng, Chunyuan, et al.
Veröffentlicht: (2026)
von: Deng, Chunyuan, et al.
Veröffentlicht: (2026)
ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
von: Anand, Nikhil, et al.
Veröffentlicht: (2026)
von: Anand, Nikhil, et al.
Veröffentlicht: (2026)
Multi-property Steering of Large Language Models with Dynamic Activation Composition
von: Scalena, Daniel, et al.
Veröffentlicht: (2024)
von: Scalena, Daniel, et al.
Veröffentlicht: (2024)
HyperSteer: Activation Steering at Scale with Hypernetworks
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
ProgGen: Generating Named Entity Recognition Datasets Step-by-step with Self-Reflexive Large Language Models
von: Heng, Yuzhao, et al.
Veröffentlicht: (2024)
von: Heng, Yuzhao, et al.
Veröffentlicht: (2024)
Steering Language Models with Weight Arithmetic
von: Fierro, Constanza, et al.
Veröffentlicht: (2025)
von: Fierro, Constanza, et al.
Veröffentlicht: (2025)
Steer Like the LLM: Activation Steering that Mimics Prompting
von: Heyman, Geert, et al.
Veröffentlicht: (2026)
von: Heyman, Geert, et al.
Veröffentlicht: (2026)
Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization
von: Cao, Yuanpu, et al.
Veröffentlicht: (2024)
von: Cao, Yuanpu, et al.
Veröffentlicht: (2024)
ROAST: Rollout-based On-distribution Activation Steering Technique
von: Su, Xuanbo, et al.
Veröffentlicht: (2026)
von: Su, Xuanbo, et al.
Veröffentlicht: (2026)
The Information Geometry of Softmax: Probing and Steering
von: Park, Kiho, et al.
Veröffentlicht: (2026)
von: Park, Kiho, et al.
Veröffentlicht: (2026)
Compositional Steering of Large Language Models with Steering Tokens
von: Radevski, Gorjan, et al.
Veröffentlicht: (2026)
von: Radevski, Gorjan, et al.
Veröffentlicht: (2026)
Guiding Giants: Lightweight Controllers for Weighted Activation Steering in LLMs
von: Hegazy, Amr, et al.
Veröffentlicht: (2025)
von: Hegazy, Amr, et al.
Veröffentlicht: (2025)
Steering MoE LLMs via Expert (De)Activation
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
SAFR: Neuron Redistribution for Interpretability
von: Chang, Ruidi, et al.
Veröffentlicht: (2025)
von: Chang, Ruidi, et al.
Veröffentlicht: (2025)
Core Context Aware Transformers for Long Context Language Modeling
von: Chen, Yaofo, et al.
Veröffentlicht: (2024)
von: Chen, Yaofo, et al.
Veröffentlicht: (2024)
Differentially Private Steering for Large Language Model Alignment
von: Goel, Anmol, et al.
Veröffentlicht: (2025)
von: Goel, Anmol, et al.
Veröffentlicht: (2025)
Programming Refusal with Conditional Activation Steering
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
SAKE: Steering Activations for Knowledge Editing
von: Scialanga, Marco, et al.
Veröffentlicht: (2025)
von: Scialanga, Marco, et al.
Veröffentlicht: (2025)
Geometry-Aware Decoding with Wasserstein-Regularized Truncation and Mass Penalties for Large Language Models
von: Davoodi, Arash Gholami, et al.
Veröffentlicht: (2026)
von: Davoodi, Arash Gholami, et al.
Veröffentlicht: (2026)
Understanding How CodeLLMs (Mis)Predict Types with Activation Steering
von: Lucchetti, Francesca, et al.
Veröffentlicht: (2024)
von: Lucchetti, Francesca, et al.
Veröffentlicht: (2024)
Massive Activations in Large Language Models
von: Sun, Mingjie, et al.
Veröffentlicht: (2024)
von: Sun, Mingjie, et al.
Veröffentlicht: (2024)
Geometry-Calibrated Conformal Abstention for Language Models
von: Xu, Rui, et al.
Veröffentlicht: (2026)
von: Xu, Rui, et al.
Veröffentlicht: (2026)
QAM-W: Joint 2D Codebook Quantization for LLM Weights via Hadamard Rotation and Activation-Aware Scaling
von: Sharma, Preetam, et al.
Veröffentlicht: (2026)
von: Sharma, Preetam, et al.
Veröffentlicht: (2026)
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
von: Laptev, Daniil, et al.
Veröffentlicht: (2025)
von: Laptev, Daniil, et al.
Veröffentlicht: (2025)
Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection
von: Deshpande, Vijeta, et al.
Veröffentlicht: (2026)
von: Deshpande, Vijeta, et al.
Veröffentlicht: (2026)
Word Embeddings Are Steers for Language Models
von: Han, Chi, et al.
Veröffentlicht: (2023)
von: Han, Chi, et al.
Veröffentlicht: (2023)
Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models
von: Li, Zihao, et al.
Veröffentlicht: (2025)
von: Li, Zihao, et al.
Veröffentlicht: (2025)
Extracting Unlearned Information from LLMs with Activation Steering
von: Seyitoğlu, Atakan, et al.
Veröffentlicht: (2024)
von: Seyitoğlu, Atakan, et al.
Veröffentlicht: (2024)
Steering Llama 2 via Contrastive Activation Addition
von: Panickssery, Nina, et al.
Veröffentlicht: (2023)
von: Panickssery, Nina, et al.
Veröffentlicht: (2023)
Activation Steering via Generative Causal Mediation
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2026)
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Language Models are Symbolic Learners in Arithmetic
von: Deng, Chunyuan, et al.
Veröffentlicht: (2024) -
Steering Information Utility in Key-Value Memory for Language Model Post-Training
von: Deng, Chunyuan, et al.
Veröffentlicht: (2025) -
Steering Language Models With Activation Engineering
von: Turner, Alexander Matt, et al.
Veröffentlicht: (2023) -
Learning Distribution-Wise Control in Representation Space for Language Models
von: Deng, Chunyuan, et al.
Veröffentlicht: (2025) -
FishBack: Pullback Fisher Geometry for Optimal Activation Steering in Transformers
von: Wang, Sihan, et al.
Veröffentlicht: (2026)