Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Pres, Itamar, Ruis, Laura, Lubana, Ekdeep Singh, Krueger, David |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
par: Jaipersaud, Brandon, et autres
Publié: (2025)
par: Jaipersaud, Brandon, et autres
Publié: (2025)
Competition Dynamics Shape Algorithmic Phases of In-Context Learning
par: Park, Core Francisco, et autres
Publié: (2024)
par: Park, Core Francisco, et autres
Publié: (2024)
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
par: Zur, Amir, et autres
Publié: (2025)
par: Zur, Amir, et autres
Publié: (2025)
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
par: Bigelow, Eric, et autres
Publié: (2025)
par: Bigelow, Eric, et autres
Publié: (2025)
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
par: Mueller, Aaron, et autres
Publié: (2025)
par: Mueller, Aaron, et autres
Publié: (2025)
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
par: Lee, Andrew, et autres
Publié: (2024)
par: Lee, Andrew, et autres
Publié: (2024)
In-Context Learning Dynamics with Random Binary Sequences
par: Bigelow, Eric J., et autres
Publié: (2023)
par: Bigelow, Eric J., et autres
Publié: (2023)
Emergence of Hierarchical Emotion Organization in Large Language Models
par: Zhao, Bo, et autres
Publié: (2025)
par: Zhao, Bo, et autres
Publié: (2025)
Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
par: Banayeeanzade, Amin, et autres
Publié: (2025)
par: Banayeeanzade, Amin, et autres
Publié: (2025)
ICLR: In-Context Learning of Representations
par: Park, Core Francisco, et autres
Publié: (2024)
par: Park, Core Francisco, et autres
Publié: (2024)
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
par: Bigelow, Eric, et autres
Publié: (2026)
par: Bigelow, Eric, et autres
Publié: (2026)
Steering Towards Fairness: Mitigating Political Bias in LLMs
par: Nadeem, Afrozah, et autres
Publié: (2025)
par: Nadeem, Afrozah, et autres
Publié: (2025)
The Depth Ceiling: On the Limits of Large Language Models in Discovering Latent Planning
par: Xu, Yi, et autres
Publié: (2026)
par: Xu, Yi, et autres
Publié: (2026)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
par: Siu, Vincent, et autres
Publié: (2025)
par: Siu, Vincent, et autres
Publié: (2025)
LiveCLKTBench: Towards Reliable Evaluation of Cross-Lingual Knowledge Transfer in Multilingual LLMs
par: Guo, Pei-Fu, et autres
Publié: (2025)
par: Guo, Pei-Fu, et autres
Publié: (2025)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
par: Bohacek, Matyas, et autres
Publié: (2025)
par: Bohacek, Matyas, et autres
Publié: (2025)
Bias Beyond Borders: Political Ideology Evaluation and Steering in Multilingual LLMs
par: Nadeem, Afrozah, et autres
Publié: (2026)
par: Nadeem, Afrozah, et autres
Publié: (2026)
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
par: Lunardi, Riccardo, et autres
Publié: (2025)
par: Lunardi, Riccardo, et autres
Publié: (2025)
Programming by Backprop: An Instruction is Worth 100 Examples When Finetuning LLMs
par: Cook, Jonathan, et autres
Publié: (2025)
par: Cook, Jonathan, et autres
Publié: (2025)
Investigating Non-Transitivity in LLM-as-a-Judge
par: Xu, Yi, et autres
Publié: (2025)
par: Xu, Yi, et autres
Publié: (2025)
CogSteer: Cognition-Inspired Selective Layer Intervention for Efficiently Steering Large Language Models
par: Wang, Xintong, et autres
Publié: (2024)
par: Wang, Xintong, et autres
Publié: (2024)
Debating with More Persuasive LLMs Leads to More Truthful Answers
par: Khan, Akbir, et autres
Publié: (2024)
par: Khan, Akbir, et autres
Publié: (2024)
Steering Language Models Before They Speak: Logit-Level Interventions
par: An, Hyeseon, et autres
Publié: (2026)
par: An, Hyeseon, et autres
Publié: (2026)
How Reliable Are Automatic Evaluation Methods for Instruction-Tuned LLMs?
par: Doostmohammadi, Ehsan, et autres
Publié: (2024)
par: Doostmohammadi, Ehsan, et autres
Publié: (2024)
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
par: Piot, Paloma, et autres
Publié: (2025)
par: Piot, Paloma, et autres
Publié: (2025)
KV Cache Steering for Controlling Frozen LLMs
par: Belitsky, Max, et autres
Publié: (2025)
par: Belitsky, Max, et autres
Publié: (2025)
Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science Communicators
par: Bajpai, Prasoon, et autres
Publié: (2024)
par: Bajpai, Prasoon, et autres
Publié: (2024)
XplainLLM: A Knowledge-Augmented Dataset for Reliable Grounded Explanations in LLMs
par: Chen, Zichen, et autres
Publié: (2023)
par: Chen, Zichen, et autres
Publié: (2023)
Towards Reliable Machine Translation: Scaling LLMs for Critical Error Detection and Safety
par: Chopra, Muskaan, et autres
Publié: (2026)
par: Chopra, Muskaan, et autres
Publié: (2026)
Double-Calibration: Towards Reliable LLMs via Calibrating Knowledge and Reasoning Confidence
par: Lu, Yuyin, et autres
Publié: (2026)
par: Lu, Yuyin, et autres
Publié: (2026)
AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production
par: Kartik, NVJK, et autres
Publié: (2025)
par: Kartik, NVJK, et autres
Publié: (2025)
Steering Awareness: Detecting Activation Steering from Within
par: Rivera, Joshua Fonseca, et autres
Publié: (2025)
par: Rivera, Joshua Fonseca, et autres
Publié: (2025)
Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions
par: Kang, Diancheng, et autres
Publié: (2026)
par: Kang, Diancheng, et autres
Publié: (2026)
Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games
par: Yadav, Neemesh, et autres
Publié: (2025)
par: Yadav, Neemesh, et autres
Publié: (2025)
Leveraging Implicit Sentiments: Enhancing Reliability and Validity in Psychological Trait Evaluation of LLMs
par: Ma, Huanhuan, et autres
Publié: (2025)
par: Ma, Huanhuan, et autres
Publié: (2025)
Exploring the Personality Traits of LLMs through Latent Features Steering
par: Yang, Shu, et autres
Publié: (2024)
par: Yang, Shu, et autres
Publié: (2024)
Probing and Steering Evaluation Awareness of Language Models
par: Nguyen, Jord, et autres
Publié: (2025)
par: Nguyen, Jord, et autres
Publié: (2025)
Characterizing and Evaluating the Reliability of LLMs against Jailbreak Attacks
par: Chen, Kexin, et autres
Publié: (2024)
par: Chen, Kexin, et autres
Publié: (2024)
Steering LLMs for Formal Theorem Proving
par: Kirtania, Shashank, et autres
Publié: (2025)
par: Kirtania, Shashank, et autres
Publié: (2025)
Steering Risk Preferences in Large Language Models by Aligning Behavioral and Neural Representations
par: Zhu, Jian-Qiao, et autres
Publié: (2025)
par: Zhu, Jian-Qiao, et autres
Publié: (2025)
Documents similaires
-
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
par: Jaipersaud, Brandon, et autres
Publié: (2025) -
Competition Dynamics Shape Algorithmic Phases of In-Context Learning
par: Park, Core Francisco, et autres
Publié: (2024) -
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
par: Zur, Amir, et autres
Publié: (2025) -
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
par: Bigelow, Eric, et autres
Publié: (2025) -
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
par: Mueller, Aaron, et autres
Publié: (2025)