Improved Representation Steering for Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Zhengxuan, Yu, Qinan, Arora, Aryaman, Manning, Christopher D., Potts, Christopher |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ReFT: Representation Finetuning for Language Models
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024)
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
di: Wu, Zhengxuan, et al.
Pubblicazione: (2025)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2025)
pyvene: A Library for Understanding and Improving PyTorch Models via Interventions
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024)
ReCOGS: How Incidental Details of a Logical Form Overshadow an Evaluation of Semantic Interpretation
di: Wu, Zhengxuan, et al.
Pubblicazione: (2023)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2023)
MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions
di: Zhong, Zexuan, et al.
Pubblicazione: (2023)
di: Zhong, Zexuan, et al.
Pubblicazione: (2023)
Language Model Circuits Are Sparse in the Neuron Basis
di: Arora, Aryaman, et al.
Pubblicazione: (2026)
di: Arora, Aryaman, et al.
Pubblicazione: (2026)
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
di: Arora, Aryaman, et al.
Pubblicazione: (2024)
di: Arora, Aryaman, et al.
Pubblicazione: (2024)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
di: Huang, Jing, et al.
Pubblicazione: (2024)
di: Huang, Jing, et al.
Pubblicazione: (2024)
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024)
HyperSteer: Activation Steering at Scale with Hypernetworks
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
PreFT: Prefill-only finetuning for efficient inference
di: Lanpouthakoun, Andrew, et al.
Pubblicazione: (2026)
di: Lanpouthakoun, Andrew, et al.
Pubblicazione: (2026)
ADAG: Automatically Describing Attribution Graphs
di: Arora, Aryaman, et al.
Pubblicazione: (2026)
di: Arora, Aryaman, et al.
Pubblicazione: (2026)
Bayesian scaling laws for in-context learning
di: Arora, Aryaman, et al.
Pubblicazione: (2024)
di: Arora, Aryaman, et al.
Pubblicazione: (2024)
Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning
di: Yu, Qinan, et al.
Pubblicazione: (2026)
di: Yu, Qinan, et al.
Pubblicazione: (2026)
Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
di: Wu, Zhengxuan, et al.
Pubblicazione: (2023)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2023)
MrT5: Dynamic Token Merging for Efficient Byte-level Language Models
di: Kallini, Julie, et al.
Pubblicazione: (2024)
di: Kallini, Julie, et al.
Pubblicazione: (2024)
Transcribe, Translate, or Transliterate: An Investigation of Intermediate Representations in Spoken Language Models
di: Ògúnrèmí, Tolúlopé, et al.
Pubblicazione: (2025)
di: Ògúnrèmí, Tolúlopé, et al.
Pubblicazione: (2025)
Mechanistic evaluation of Transformers and state space models
di: Arora, Aryaman, et al.
Pubblicazione: (2025)
di: Arora, Aryaman, et al.
Pubblicazione: (2025)
Sneaking Syntax into Transformer Language Models with Tree Regularization
di: Nandi, Ananjan, et al.
Pubblicazione: (2024)
di: Nandi, Ananjan, et al.
Pubblicazione: (2024)
Improving Multilingual Language Models by Aligning Representations through Steering
di: Mahmoud, Omar, et al.
Pubblicazione: (2025)
di: Mahmoud, Omar, et al.
Pubblicazione: (2025)
Drop Dropout on Single-Epoch Language Model Pretraining
di: Liu, Houjun, et al.
Pubblicazione: (2025)
di: Liu, Houjun, et al.
Pubblicazione: (2025)
Stronger Baselines for Retrieval-Augmented Generation with Long-Context Language Models
di: Laitenberger, Alex, et al.
Pubblicazione: (2025)
di: Laitenberger, Alex, et al.
Pubblicazione: (2025)
CGELBank Annotation Manual v1.2
di: Reynolds, Brett, et al.
Pubblicazione: (2023)
di: Reynolds, Brett, et al.
Pubblicazione: (2023)
Language models as tools for investigating the distinction between possible and impossible natural languages
di: Kallini, Julie, et al.
Pubblicazione: (2025)
di: Kallini, Julie, et al.
Pubblicazione: (2025)
Do Language Models Use Their Depth Efficiently?
di: Csordás, Róbert, et al.
Pubblicazione: (2025)
di: Csordás, Róbert, et al.
Pubblicazione: (2025)
Base Models Beat Aligned Models at Randomness and Creativity
di: West, Peter, et al.
Pubblicazione: (2025)
di: West, Peter, et al.
Pubblicazione: (2025)
On the Limitations of Steering in Language Model Alignment
di: Niranjan, Chebrolu, et al.
Pubblicazione: (2025)
di: Niranjan, Chebrolu, et al.
Pubblicazione: (2025)
Humans and transformer LMs: Abstraction drives language learning
di: Jian, Jasper, et al.
Pubblicazione: (2026)
di: Jian, Jasper, et al.
Pubblicazione: (2026)
Demystifying Verbatim Memorization in Large Language Models
di: Huang, Jing, et al.
Pubblicazione: (2024)
di: Huang, Jing, et al.
Pubblicazione: (2024)
The Cylindrical Representation Hypothesis for Language Model Steering
di: Gao, Lang, et al.
Pubblicazione: (2026)
di: Gao, Lang, et al.
Pubblicazione: (2026)
Improving Pretraining Data Using Perplexity Correlations
di: Thrush, Tristan, et al.
Pubblicazione: (2024)
di: Thrush, Tristan, et al.
Pubblicazione: (2024)
False Friends Are Not Foes: Investigating Vocabulary Overlap in Multilingual Language Models
di: Kallini, Julie, et al.
Pubblicazione: (2025)
di: Kallini, Julie, et al.
Pubblicazione: (2025)
A paradox of AI fluency
di: Potts, Christopher, et al.
Pubblicazione: (2026)
di: Potts, Christopher, et al.
Pubblicazione: (2026)
Invisible failures in human-AI interactions
di: Potts, Christopher, et al.
Pubblicazione: (2026)
di: Potts, Christopher, et al.
Pubblicazione: (2026)
Osiris: A Lightweight Open-Source Hallucination Detection System
di: Shan, Alex, et al.
Pubblicazione: (2025)
di: Shan, Alex, et al.
Pubblicazione: (2025)
Towards Inference-time Category-wise Safety Steering for Large Language Models
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2024)
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2024)
Self-Improving Model Steering
di: Zhu, Rongyi, et al.
Pubblicazione: (2025)
di: Zhu, Rongyi, et al.
Pubblicazione: (2025)
Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
di: Csordás, Róbert, et al.
Pubblicazione: (2024)
di: Csordás, Róbert, et al.
Pubblicazione: (2024)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
BAGEL: Bootstrapping Agents by Guiding Exploration with Language
di: Murty, Shikhar, et al.
Pubblicazione: (2024)
di: Murty, Shikhar, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ReFT: Representation Finetuning for Language Models
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024) -
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
di: Wu, Zhengxuan, et al.
Pubblicazione: (2025) -
pyvene: A Library for Understanding and Improving PyTorch Models via Interventions
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024) -
ReCOGS: How Incidental Details of a Logical Form Overshadow an Evaluation of Semantic Interpretation
di: Wu, Zhengxuan, et al.
Pubblicazione: (2023) -
MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions
di: Zhong, Zexuan, et al.
Pubblicazione: (2023)