The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness
Fuente:
arXiv
Saved in:
| Main Authors: | Abdelnabi, Sahar, Salem, Ahmed |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
by: Gomaa, Amr, et al.
Published: (2025)
by: Gomaa, Amr, et al.
Published: (2025)
AI Agents May Always Fall for Prompt Injections
by: Abdelnabi, Sahar, et al.
Published: (2026)
by: Abdelnabi, Sahar, et al.
Published: (2026)
Get my drift? Catching LLM Task Drift with Activation Deltas
by: Abdelnabi, Sahar, et al.
Published: (2024)
by: Abdelnabi, Sahar, et al.
Published: (2024)
Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation
by: Abdelnabi, Sahar, et al.
Published: (2023)
by: Abdelnabi, Sahar, et al.
Published: (2023)
Evaluation Awareness in Language Models Has Limited Effect on Behaviour
by: Knecht, Amelie, et al.
Published: (2026)
by: Knecht, Amelie, et al.
Published: (2026)
QSTN: A Modular Framework for Robust Questionnaire Inference with Large Language Models
by: Kreutner, Maximilian, et al.
Published: (2025)
by: Kreutner, Maximilian, et al.
Published: (2025)
Models That Know How Evaluations Are Designed Score Safer
by: Deckenbach, Katharina, et al.
Published: (2026)
by: Deckenbach, Katharina, et al.
Published: (2026)
Evaluating Proactive Risk Awareness of Large Language Models
by: Luo, Xuan, et al.
Published: (2026)
by: Luo, Xuan, et al.
Published: (2026)
Modeling Motivated Reasoning in Law: Evaluating Strategic Role Conditioning in LLM Summarization
by: Cho, Eunjung, et al.
Published: (2025)
by: Cho, Eunjung, et al.
Published: (2025)
Decomposing and Measuring Evaluation Awareness
by: Li, Changling, et al.
Published: (2026)
by: Li, Changling, et al.
Published: (2026)
MalAlgoQA: Pedagogical Evaluation of Counterfactual Reasoning in Large Language Models and Implications for AI in Education
by: Liu, Naiming, et al.
Published: (2024)
by: Liu, Naiming, et al.
Published: (2024)
Toward Socially Aware Vision-Language Models: Evaluating Cultural Competence Through Multimodal Story Generation
by: Mukherjee, Arka, et al.
Published: (2025)
by: Mukherjee, Arka, et al.
Published: (2025)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
by: Kabir, Mohsinul, et al.
Published: (2026)
by: Kabir, Mohsinul, et al.
Published: (2026)
PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning
by: Akyürek, Afra Feyza, et al.
Published: (2025)
by: Akyürek, Afra Feyza, et al.
Published: (2025)
Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning
by: Juvekar, Kush, et al.
Published: (2025)
by: Juvekar, Kush, et al.
Published: (2025)
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
by: Oh, Gyutaek, et al.
Published: (2025)
by: Oh, Gyutaek, et al.
Published: (2025)
DaKultur: Evaluating the Cultural Awareness of Language Models for Danish with Native Speakers
by: Müller-Eberstein, Max, et al.
Published: (2025)
by: Müller-Eberstein, Max, et al.
Published: (2025)
DSO: Direct Steering Optimization for Bias Mitigation
by: Paes, Lucas Monteiro, et al.
Published: (2025)
by: Paes, Lucas Monteiro, et al.
Published: (2025)
White-Box Sensitivity Auditing with Steering Vectors
by: Cyberey, Hannah, et al.
Published: (2026)
by: Cyberey, Hannah, et al.
Published: (2026)
Representation Surgery: Theory and Practice of Affine Steering
by: Singh, Shashwat, et al.
Published: (2024)
by: Singh, Shashwat, et al.
Published: (2024)
Assessing Judging Bias in Large Reasoning Models: An Empirical Study
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
by: Ghosh, Rajarshi, et al.
Published: (2025)
by: Ghosh, Rajarshi, et al.
Published: (2025)
Steering at the Source: Style Modulation Heads for Robust Persona Control
by: Izawa, Yoshihiro, et al.
Published: (2026)
by: Izawa, Yoshihiro, et al.
Published: (2026)
JudgeMeNot: Personalizing Large Language Models to Emulate Judicial Reasoning in Hebrew
by: Razumenko, Itay, et al.
Published: (2026)
by: Razumenko, Itay, et al.
Published: (2026)
Evaluating Cultural Awareness of LLMs for Yoruba, Malayalam, and English
by: Dawson, Fiifi, et al.
Published: (2024)
by: Dawson, Fiifi, et al.
Published: (2024)
A Theory of Response Sampling in LLMs: Part Descriptive and Part Prescriptive
by: Sivaprasad, Sarath, et al.
Published: (2024)
by: Sivaprasad, Sarath, et al.
Published: (2024)
Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
by: Cyberey, Hannah, et al.
Published: (2025)
by: Cyberey, Hannah, et al.
Published: (2025)
Translate With Care: Addressing Gender Bias, Neutrality, and Reasoning in Large Language Model Translations
by: Zahraei, Pardis Sadat, et al.
Published: (2025)
by: Zahraei, Pardis Sadat, et al.
Published: (2025)
Inference-Time Reasoning Selectively Reduces Implicit Social Bias in Large Language Models
by: Apsel, Molly, et al.
Published: (2026)
by: Apsel, Molly, et al.
Published: (2026)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
by: Tang, Zeyu, et al.
Published: (2026)
by: Tang, Zeyu, et al.
Published: (2026)
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
by: Sahoo, Subramanyam, et al.
Published: (2026)
by: Sahoo, Subramanyam, et al.
Published: (2026)
Taxonomy, Opportunities, and Challenges of Representation Engineering for Large Language Models
by: Wehner, Jan, et al.
Published: (2025)
by: Wehner, Jan, et al.
Published: (2025)
Reasoning Language Models for complex assessments tasks: Evaluating parental cooperation from child protection case reports
by: Stoll, Dragan, et al.
Published: (2026)
by: Stoll, Dragan, et al.
Published: (2026)
Extrinsic Evaluation of Cultural Competence in Large Language Models
by: Bhatt, Shaily, et al.
Published: (2024)
by: Bhatt, Shaily, et al.
Published: (2024)
Out of the Box, into the Clinic? Evaluating State-of-the-Art ASR for Clinical Applications for Older Adults
by: van Dijk, Bram, et al.
Published: (2025)
by: van Dijk, Bram, et al.
Published: (2025)
Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models
by: Choi, Younwoo, et al.
Published: (2025)
by: Choi, Younwoo, et al.
Published: (2025)
Narrative over Numbers: The Identifiable Victim Effect and its Amplification Under Alignment and Reasoning in Large Language Models
by: Raiyan, Syed Rifat
Published: (2026)
by: Raiyan, Syed Rifat
Published: (2026)
Firewalls to Secure Dynamic LLM Agentic Networks
by: Abdelnabi, Sahar, et al.
Published: (2025)
by: Abdelnabi, Sahar, et al.
Published: (2025)
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
by: Jahara, Fatima, et al.
Published: (2025)
by: Jahara, Fatima, et al.
Published: (2025)
RDBE: Reasoning Distillation-Based Evaluation Enhances Automatic Essay Scoring
by: Mohammadkhani, Ali Ghiasvand
Published: (2024)
by: Mohammadkhani, Ali Ghiasvand
Published: (2024)
Similar Items
-
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
by: Gomaa, Amr, et al.
Published: (2025) -
AI Agents May Always Fall for Prompt Injections
by: Abdelnabi, Sahar, et al.
Published: (2026) -
Get my drift? Catching LLM Task Drift with Activation Deltas
by: Abdelnabi, Sahar, et al.
Published: (2024) -
Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation
by: Abdelnabi, Sahar, et al.
Published: (2023) -
Evaluation Awareness in Language Models Has Limited Effect on Behaviour
by: Knecht, Amelie, et al.
Published: (2026)