Evaluation Awareness Scales Predictably in Open-Weights Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909893775065088 |
|---|---|
| author | Chaudhary, Maheep Su, Ian Hooda, Nikhil Shankar, Nishith Tan, Julia Zhu, Kevin Lagasse, Ryan Sharma, Vasu Panda, Ashwinee |
| author_facet | Chaudhary, Maheep Su, Ian Hooda, Nikhil Shankar, Nishith Tan, Julia Zhu, Kevin Lagasse, Ryan Sharma, Vasu Panda, Ashwinee |
| contents | Large language models (LLMs) can internally distinguish between evaluation and deployment contexts, a behaviour known as \emph{evaluation awareness}. This undermines AI safety evaluations, as models may conceal dangerous capabilities during testing. Prior work demonstrated this in a single $70$B model, but the scaling relationship across model sizes remains unknown. We investigate evaluation awareness across $15$ models scaling from $0.27$B to $70$B parameters from four families using linear probing on steering vector activations. Our results reveal a clear power-law scaling: evaluation awareness increases predictably with model size. This scaling law enables forecasting deceptive behavior in future larger models and guides the design of scale-aware evaluation strategies for AI safety. A link to the implementation of this paper can be found at https://anonymous.4open.science/r/evaluation-awareness-scaling-laws/README.md. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_13333 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Evaluation Awareness Scales Predictably in Open-Weights Large Language Models Chaudhary, Maheep Su, Ian Hooda, Nikhil Shankar, Nishith Tan, Julia Zhu, Kevin Lagasse, Ryan Sharma, Vasu Panda, Ashwinee Artificial Intelligence Large language models (LLMs) can internally distinguish between evaluation and deployment contexts, a behaviour known as \emph{evaluation awareness}. This undermines AI safety evaluations, as models may conceal dangerous capabilities during testing. Prior work demonstrated this in a single $70$B model, but the scaling relationship across model sizes remains unknown. We investigate evaluation awareness across $15$ models scaling from $0.27$B to $70$B parameters from four families using linear probing on steering vector activations. Our results reveal a clear power-law scaling: evaluation awareness increases predictably with model size. This scaling law enables forecasting deceptive behavior in future larger models and guides the design of scale-aware evaluation strategies for AI safety. A link to the implementation of this paper can be found at https://anonymous.4open.science/r/evaluation-awareness-scaling-laws/README.md. |
| title | Evaluation Awareness Scales Predictably in Open-Weights Large Language Models |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2509.13333 |