Steering Evaluation-Aware Language Models to Act Like They Are Deployed
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hua, Tim Tian, Qin, Andrew, Marks, Samuel, Nanda, Neel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
von: Casademunt, Helena, et al.
Veröffentlicht: (2025)
von: Casademunt, Helena, et al.
Veröffentlicht: (2025)
Explorations of Self-Repair in Language Models
von: Rushing, Cody, et al.
Veröffentlicht: (2024)
von: Rushing, Cody, et al.
Veröffentlicht: (2024)
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
von: Zhang, Fred, et al.
Veröffentlicht: (2023)
von: Zhang, Fred, et al.
Veröffentlicht: (2023)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
von: Ferrando, Javier, et al.
Veröffentlicht: (2024)
von: Ferrando, Javier, et al.
Veröffentlicht: (2024)
Probing and Steering Evaluation Awareness of Language Models
von: Nguyen, Jord, et al.
Veröffentlicht: (2025)
von: Nguyen, Jord, et al.
Veröffentlicht: (2025)
Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation
von: Casademunt, Helena, et al.
Veröffentlicht: (2026)
von: Casademunt, Helena, et al.
Veröffentlicht: (2026)
Liars' Bench: Evaluating Lie Detectors for Language Models
von: Kretschmar, Kieron, et al.
Veröffentlicht: (2025)
von: Kretschmar, Kieron, et al.
Veröffentlicht: (2025)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
von: Soligo, Anna, et al.
Veröffentlicht: (2026)
von: Soligo, Anna, et al.
Veröffentlicht: (2026)
Steering Large Language Models to Evaluate and Amplify Creativity
von: Olson, Matthew Lyle, et al.
Veröffentlicht: (2024)
von: Olson, Matthew Lyle, et al.
Veröffentlicht: (2024)
Steering Multimodal Large Language Models Decoding for Context-Aware Safety
von: Liu, Zheyuan, et al.
Veröffentlicht: (2025)
von: Liu, Zheyuan, et al.
Veröffentlicht: (2025)
Confidence Regulation Neurons in Language Models
von: Stolfo, Alessandro, et al.
Veröffentlicht: (2024)
von: Stolfo, Alessandro, et al.
Veröffentlicht: (2024)
Refusal in Language Models Is Mediated by a Single Direction
von: Arditi, Andy, et al.
Veröffentlicht: (2024)
von: Arditi, Andy, et al.
Veröffentlicht: (2024)
Steering Awareness: Detecting Activation Steering from Within
von: Rivera, Joshua Fonseca, et al.
Veröffentlicht: (2025)
von: Rivera, Joshua Fonseca, et al.
Veröffentlicht: (2025)
Self-Steering Language Models
von: Grand, Gabriel, et al.
Veröffentlicht: (2025)
von: Grand, Gabriel, et al.
Veröffentlicht: (2025)
Steer Like the LLM: Activation Steering that Mimics Prompting
von: Heyman, Geert, et al.
Veröffentlicht: (2026)
von: Heyman, Geert, et al.
Veröffentlicht: (2026)
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
von: Karvonen, Adam, et al.
Veröffentlicht: (2024)
von: Karvonen, Adam, et al.
Veröffentlicht: (2024)
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
von: Mekky, Ali, et al.
Veröffentlicht: (2025)
von: Mekky, Ali, et al.
Veröffentlicht: (2025)
Steering When Necessary: Flexible Steering Large Language Models with Backtracking
von: Cheng, Zifeng, et al.
Veröffentlicht: (2025)
von: Cheng, Zifeng, et al.
Veröffentlicht: (2025)
LLMCBench: Benchmarking Large Language Model Compression for Efficient Deployment
von: Yang, Ge, et al.
Veröffentlicht: (2024)
von: Yang, Ge, et al.
Veröffentlicht: (2024)
On the Limitations of Steering in Language Model Alignment
von: Niranjan, Chebrolu, et al.
Veröffentlicht: (2025)
von: Niranjan, Chebrolu, et al.
Veröffentlicht: (2025)
Towards Evaluating Proactive Risk Awareness of Multimodal Language Models
von: Yuan, Youliang, et al.
Veröffentlicht: (2025)
von: Yuan, Youliang, et al.
Veröffentlicht: (2025)
Universal Neurons in GPT2 Language Models
von: Gurnee, Wes, et al.
Veröffentlicht: (2024)
von: Gurnee, Wes, et al.
Veröffentlicht: (2024)
Internal states before wait modulate reasoning patterns
von: Troitskii, Dmitrii, et al.
Veröffentlicht: (2025)
von: Troitskii, Dmitrii, et al.
Veröffentlicht: (2025)
BILLY: Steering Large Language Models via Merging Persona Vectors for Creative Generation
von: Pai, Tsung-Min, et al.
Veröffentlicht: (2025)
von: Pai, Tsung-Min, et al.
Veröffentlicht: (2025)
Activation Scaling for Steering and Interpreting Language Models
von: Stoehr, Niklas, et al.
Veröffentlicht: (2024)
von: Stoehr, Niklas, et al.
Veröffentlicht: (2024)
Thought Branches: Interpreting LLM Reasoning Requires Resampling
von: Macar, Uzay, et al.
Veröffentlicht: (2025)
von: Macar, Uzay, et al.
Veröffentlicht: (2025)
Thought Anchors: Which LLM Reasoning Steps Matter?
von: Bogdan, Paul C., et al.
Veröffentlicht: (2025)
von: Bogdan, Paul C., et al.
Veröffentlicht: (2025)
FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models
von: Weng, Zixuan, et al.
Veröffentlicht: (2026)
von: Weng, Zixuan, et al.
Veröffentlicht: (2026)
CogSteer: Cognition-Inspired Selective Layer Intervention for Efficiently Steering Large Language Models
von: Wang, Xintong, et al.
Veröffentlicht: (2024)
von: Wang, Xintong, et al.
Veröffentlicht: (2024)
Compositional Steering of Large Language Models with Steering Tokens
von: Radevski, Gorjan, et al.
Veröffentlicht: (2026)
von: Radevski, Gorjan, et al.
Veröffentlicht: (2026)
Prompt-Based Value Steering of Large Language Models
von: Abbo, Giulio Antonio, et al.
Veröffentlicht: (2025)
von: Abbo, Giulio Antonio, et al.
Veröffentlicht: (2025)
Cross-Lingual Activation Steering for Multilingual Language Models
von: Pokharel, Rhitabrat, et al.
Veröffentlicht: (2026)
von: Pokharel, Rhitabrat, et al.
Veröffentlicht: (2026)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
von: Soo, Samuel, et al.
Veröffentlicht: (2025)
von: Soo, Samuel, et al.
Veröffentlicht: (2025)
NutriBench: A Dataset for Evaluating Large Language Models on Nutrition Estimation from Meal Descriptions
von: Hua, Andong, et al.
Veröffentlicht: (2024)
von: Hua, Andong, et al.
Veröffentlicht: (2024)
Scaling sparse feature circuit finding for in-context learning
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
von: Minder, Julian, et al.
Veröffentlicht: (2025)
von: Minder, Julian, et al.
Veröffentlicht: (2025)
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
von: Maar, Jim, et al.
Veröffentlicht: (2026)
von: Maar, Jim, et al.
Veröffentlicht: (2026)
Robustly Improving LLM Fairness in Realistic Settings via Interpretability
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
von: Minder, Julian, et al.
Veröffentlicht: (2025)
von: Minder, Julian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
von: Casademunt, Helena, et al.
Veröffentlicht: (2025) -
Explorations of Self-Repair in Language Models
von: Rushing, Cody, et al.
Veröffentlicht: (2024) -
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
von: Zhang, Fred, et al.
Veröffentlicht: (2023) -
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
von: Ferrando, Javier, et al.
Veröffentlicht: (2024) -
Probing and Steering Evaluation Awareness of Language Models
von: Nguyen, Jord, et al.
Veröffentlicht: (2025)