Evaluation Awareness Scales Predictably in Open-Weights Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chaudhary, Maheep, Su, Ian, Hooda, Nikhil, Shankar, Nishith, Tan, Julia, Zhu, Kevin, Lagasse, Ryan, Sharma, Vasu, Panda, Ashwinee
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909893775065088
author Chaudhary, Maheep
Su, Ian
Hooda, Nikhil
Shankar, Nishith
Tan, Julia
Zhu, Kevin
Lagasse, Ryan
Sharma, Vasu
Panda, Ashwinee
author_facet Chaudhary, Maheep
Su, Ian
Hooda, Nikhil
Shankar, Nishith
Tan, Julia
Zhu, Kevin
Lagasse, Ryan
Sharma, Vasu
Panda, Ashwinee
contents Large language models (LLMs) can internally distinguish between evaluation and deployment contexts, a behaviour known as \emph{evaluation awareness}. This undermines AI safety evaluations, as models may conceal dangerous capabilities during testing. Prior work demonstrated this in a single $70$B model, but the scaling relationship across model sizes remains unknown. We investigate evaluation awareness across $15$ models scaling from $0.27$B to $70$B parameters from four families using linear probing on steering vector activations. Our results reveal a clear power-law scaling: evaluation awareness increases predictably with model size. This scaling law enables forecasting deceptive behavior in future larger models and guides the design of scale-aware evaluation strategies for AI safety. A link to the implementation of this paper can be found at https://anonymous.4open.science/r/evaluation-awareness-scaling-laws/README.md.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13333
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluation Awareness Scales Predictably in Open-Weights Large Language Models
Chaudhary, Maheep
Su, Ian
Hooda, Nikhil
Shankar, Nishith
Tan, Julia
Zhu, Kevin
Lagasse, Ryan
Sharma, Vasu
Panda, Ashwinee
Artificial Intelligence
Large language models (LLMs) can internally distinguish between evaluation and deployment contexts, a behaviour known as \emph{evaluation awareness}. This undermines AI safety evaluations, as models may conceal dangerous capabilities during testing. Prior work demonstrated this in a single $70$B model, but the scaling relationship across model sizes remains unknown. We investigate evaluation awareness across $15$ models scaling from $0.27$B to $70$B parameters from four families using linear probing on steering vector activations. Our results reveal a clear power-law scaling: evaluation awareness increases predictably with model size. This scaling law enables forecasting deceptive behavior in future larger models and guides the design of scale-aware evaluation strategies for AI safety. A link to the implementation of this paper can be found at https://anonymous.4open.science/r/evaluation-awareness-scaling-laws/README.md.
title Evaluation Awareness Scales Predictably in Open-Weights Large Language Models
topic Artificial Intelligence
url https://arxiv.org/abs/2509.13333