Evaluation Awareness Scales Predictably in Open-Weights Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Chaudhary, Maheep, Su, Ian, Hooda, Nikhil, Shankar, Nishith, Tan, Julia, Zhu, Kevin, Lagasse, Ryan, Sharma, Vasu, Panda, Ashwinee |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
by: Patel, Dev, et al.
Published: (2025)
by: Patel, Dev, et al.
Published: (2025)
SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
by: Batra, Shourya, et al.
Published: (2025)
by: Batra, Shourya, et al.
Published: (2025)
In-Context Environments Induce Evaluation-Awareness in Language Models
by: Chaudhary, Maheep
Published: (2026)
by: Chaudhary, Maheep
Published: (2026)
FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness
by: Swaroop, Anand, et al.
Published: (2025)
by: Swaroop, Anand, et al.
Published: (2025)
Amortized Latent Steering: Low-Cost Alternative to Test-Time Optimization
by: Egbuna, Nathan, et al.
Published: (2025)
by: Egbuna, Nathan, et al.
Published: (2025)
Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
by: Chaudhary, Maheep, et al.
Published: (2024)
by: Chaudhary, Maheep, et al.
Published: (2024)
Weight space Detection of Backdoors in LoRA Adapters
by: Merenciano, David Puertolas, et al.
Published: (2026)
by: Merenciano, David Puertolas, et al.
Published: (2026)
A Few Bad Neurons: Isolating and Surgically Correcting Sycophancy
by: O'Brien, Claire, et al.
Published: (2026)
by: O'Brien, Claire, et al.
Published: (2026)
Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs
by: Chaturvedi, Isha, et al.
Published: (2025)
by: Chaturvedi, Isha, et al.
Published: (2025)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
by: Chaudhary, Maheep, et al.
Published: (2025)
by: Chaudhary, Maheep, et al.
Published: (2025)
Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?
by: Nunez, Jeanmely Rojas, et al.
Published: (2026)
by: Nunez, Jeanmely Rojas, et al.
Published: (2026)
Modeling and Predicting Multi-Turn Answer Instability in Large Language Models
by: He, Jiahang, et al.
Published: (2025)
by: He, Jiahang, et al.
Published: (2025)
Broken Chains: The Cost of Incomplete Reasoning in LLMs
by: Su, Ian, et al.
Published: (2026)
by: Su, Ian, et al.
Published: (2026)
Hydra: A Modular Architecture for Efficient Long-Context Reasoning
by: Chaudhary, Siddharth, et al.
Published: (2025)
by: Chaudhary, Siddharth, et al.
Published: (2025)
@GrokSet: multi-party Human-LLM Interactions in Social Media
by: Migliarini, Matteo, et al.
Published: (2026)
by: Migliarini, Matteo, et al.
Published: (2026)
Probe-Rewrite-Evaluate: A Workflow for Reliable Benchmarks and Quantifying Evaluation Awareness
by: Xiong, Lang, et al.
Published: (2025)
by: Xiong, Lang, et al.
Published: (2025)
A New Linear Scaling Rule for Private Adaptive Hyperparameter Optimization
by: Panda, Ashwinee, et al.
Published: (2022)
by: Panda, Ashwinee, et al.
Published: (2022)
MANATEE: Inference-Time Lightweight Diffusion Based Safety Defense for LLMs
by: Kan, Chun Yan Ryan, et al.
Published: (2026)
by: Kan, Chun Yan Ryan, et al.
Published: (2026)
Interactions of NK Cells and Macrophages: From Infections to Cancer Therapeutics
by: Vishakha Hooda, et al.
Published: (2024)
by: Vishakha Hooda, et al.
Published: (2024)
FineGRAIN: Evaluating Failure Modes of Text-to-Image Models with Vision Language Model Judges
by: Hayes, Kevin David, et al.
Published: (2025)
by: Hayes, Kevin David, et al.
Published: (2025)
A Scaling Law for Token Efficiency in LLM Fine-Tuning Under Fixed Compute Budgets
by: Lagasse, Ryan, et al.
Published: (2025)
by: Lagasse, Ryan, et al.
Published: (2025)
Private Fine-tuning of Large Language Models with Zeroth-order Optimization
by: Tang, Xinyu, et al.
Published: (2024)
by: Tang, Xinyu, et al.
Published: (2024)
Optimizing Chain-of-Thought Confidence via Topological and Dirichlet Risk Analysis
by: More, Abhishek, et al.
Published: (2025)
by: More, Abhishek, et al.
Published: (2025)
Reasoning Relay: Evaluating Stability and Interchangeability of Large Language Models in Mathematical Reasoning
by: Lu, Leo, et al.
Published: (2025)
by: Lu, Leo, et al.
Published: (2025)
Multi-Token Prediction via Self-Distillation
by: Kirchenbauer, John, et al.
Published: (2026)
by: Kirchenbauer, John, et al.
Published: (2026)
LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation
by: Zhang, Juzheng, et al.
Published: (2025)
by: Zhang, Juzheng, et al.
Published: (2025)
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification
by: Shah, Siddhant Bikram, et al.
Published: (2024)
by: Shah, Siddhant Bikram, et al.
Published: (2024)
Privacy Auditing of Large Language Models
by: Panda, Ashwinee, et al.
Published: (2025)
by: Panda, Ashwinee, et al.
Published: (2025)
Shared Parameter Subspaces and Cross-Task Linearity in Emergently Misaligned Behavior
by: Arturi, Daniel Aarao Reis, et al.
Published: (2025)
by: Arturi, Daniel Aarao Reis, et al.
Published: (2025)
Discovering Transformer Circuits via a Hybrid Attribution and Pruning Framework
by: Gu, Hao, et al.
Published: (2025)
by: Gu, Hao, et al.
Published: (2025)
Rewrite-to-Rank: Optimizing Ad Visibility via Retrieval-Aware Text Rewriting
by: Ho, Chloe, et al.
Published: (2025)
by: Ho, Chloe, et al.
Published: (2025)
CIPHER: Cryptographic Insecurity Profiling via Hybrid Evaluation of Responses
by: Manolov, Max, et al.
Published: (2026)
by: Manolov, Max, et al.
Published: (2026)
Eco‐friendly Regioselective Synthesis, Biological Evaluation of Some New 5‐acylfunctionalized 2‐(1H‐pyrazol‐1‐yl)thiazoles as Potential Antimicrobial and Anthelmintic Agents
by: Ranjana Aggarwal, et al.
Published: (2024)
by: Ranjana Aggarwal, et al.
Published: (2024)
Punctuation and Predicates in Language Models
by: Chauhan, Sonakshi, et al.
Published: (2025)
by: Chauhan, Sonakshi, et al.
Published: (2025)
Studying Cross-cluster Modularity in Neural Networks
by: Golechha, Satvik, et al.
Published: (2025)
by: Golechha, Satvik, et al.
Published: (2025)
PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases
by: Vuddanti, Sri Vatsa, et al.
Published: (2025)
by: Vuddanti, Sri Vatsa, et al.
Published: (2025)
Quantifying Seasonal Weather Risk in Indian Markets: Stochastic Model for Risk-Averse State-Specific Temperature Derivative Pricing
by: Hooda, Soumil, et al.
Published: (2024)
by: Hooda, Soumil, et al.
Published: (2024)
The Rules of the Coronation: Differentiating Convention from Practice and Custom
by: Carolyn S. Harris, et al.
Published: (2025)
by: Carolyn S. Harris, et al.
Published: (2025)
Scaling Open-Ended Reasoning to Predict the Future
by: Chandak, Nikhil, et al.
Published: (2025)
by: Chandak, Nikhil, et al.
Published: (2025)
Analysis of Attention in Video Diffusion Transformers
by: Wen, Yuxin, et al.
Published: (2025)
by: Wen, Yuxin, et al.
Published: (2025)
Similar Items
-
Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
by: Patel, Dev, et al.
Published: (2025) -
SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
by: Batra, Shourya, et al.
Published: (2025) -
In-Context Environments Induce Evaluation-Awareness in Language Models
by: Chaudhary, Maheep
Published: (2026) -
FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness
by: Swaroop, Anand, et al.
Published: (2025) -
Amortized Latent Steering: Low-Cost Alternative to Test-Time Optimization
by: Egbuna, Nathan, et al.
Published: (2025)