Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Tice, Cameron, Kreer, Philipp Alexander, Helm-Burger, Nathan, Shahani, Prithviraj Singh, Ryzhenkov, Fedor, Roger, Fabien, Neo, Clement, Haimes, Jacob, Hofstätter, Felix, van der Weij, Teun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
by: van der Weij, Teun, et al.
Published: (2024)
by: van der Weij, Teun, et al.
Published: (2024)
The Elicitation Game: Evaluating Capability Elicitation Techniques
by: Hofstätter, Felix, et al.
Published: (2025)
by: Hofstätter, Felix, et al.
Published: (2025)
Noise Injection Systemically Degrades Large Language Model Safety Guardrails
by: Shahani, Prithviraj Singh, et al.
Published: (2025)
by: Shahani, Prithviraj Singh, et al.
Published: (2025)
Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
by: Haimes, Jacob, et al.
Published: (2024)
by: Haimes, Jacob, et al.
Published: (2024)
Extending Activation Steering to Broad Skills and Multiple Behaviours
by: van der Weij, Teun, et al.
Published: (2024)
by: van der Weij, Teun, et al.
Published: (2024)
Auditing Games for Sandbagging
by: Taylor, Jordan, et al.
Published: (2025)
by: Taylor, Jordan, et al.
Published: (2025)
Evaluating Precise Geolocation Inference Capabilities of Vision Language Models
by: Jay, Neel, et al.
Published: (2025)
by: Jay, Neel, et al.
Published: (2025)
LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
by: Li, Chloe, et al.
Published: (2025)
by: Li, Chloe, et al.
Published: (2025)
Option-Order Randomisation Reveals a Distributional Position Attractor in Prompted Sandbagging
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Probing a Vision-Language-Action Model for Symbolic States and Integration into a Cognitive Architecture
by: Lu, Hong, et al.
Published: (2025)
by: Lu, Hong, et al.
Published: (2025)
Chord Colourizer: A Near Real-Time System for Visualizing Musical Key
by: Haimes, Paul
Published: (2025)
by: Haimes, Paul
Published: (2025)
Tailored Truths: Optimizing LLM Persuasion with Personalization and Fabricated Statistics
by: Timm, Jasper, et al.
Published: (2025)
by: Timm, Jasper, et al.
Published: (2025)
View From Above: A Framework for Evaluating Distribution Shifts in Model Behavior
by: Chopra, Tanush, et al.
Published: (2024)
by: Chopra, Tanush, et al.
Published: (2024)
Sandbagging in a Simple Survival Bandit Problem
by: Dyer, Joel, et al.
Published: (2025)
by: Dyer, Joel, et al.
Published: (2025)
Removing Sandbagging in LLMs by Training with Weak Supervision
by: Ryd, Emil, et al.
Published: (2026)
by: Ryd, Emil, et al.
Published: (2026)
The H-graph with unequal masses in quantum field theory
by: Kreer, Philipp Alexander, et al.
Published: (2024)
by: Kreer, Philipp Alexander, et al.
Published: (2024)
The H-graph with equal masses in terms of multiple polylogarithms
by: Kreer, Philipp Alexander, et al.
Published: (2021)
by: Kreer, Philipp Alexander, et al.
Published: (2021)
EGT Quantum Resonance: Predicting the $\mathbf{402 \text{ GeV}}$ WIMP and Future Discovery Conditions at the Large Hadron Collider
by: Tice, Brian
Published: (2025)
by: Tice, Brian
Published: (2025)
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
by: Sturgeon, Benjamin, et al.
Published: (2025)
by: Sturgeon, Benjamin, et al.
Published: (2025)
Trace Relations in Deformed Gauge Theories
by: Raman, Madhusudhan, et al.
Published: (2024)
by: Raman, Madhusudhan, et al.
Published: (2024)
The Anatomy of Alignment: Decomposing Preference Optimization by Steering Sparse Features
by: Ferrao, Jeremias, et al.
Published: (2025)
by: Ferrao, Jeremias, et al.
Published: (2025)
Direct Ascent Synthesis: Revealing Hidden Generative Capabilities in Discriminative Models
by: Fort, Stanislav, et al.
Published: (2025)
by: Fort, Stanislav, et al.
Published: (2025)
Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents
by: Turk, Matt
Published: (2026)
by: Turk, Matt
Published: (2026)
Definição e Percepção de Imagem: Um Estudo em uma Escola de Educação Infantil de Novo Hamburgo
by: Cássia Rebelo Hofstätter
Published: (2008)
by: Cássia Rebelo Hofstätter
Published: (2008)
A Pesquisa de Marketing como um Meio de Informação para a Tomada de Decisão Estratégica
by: Cássia Rebelo Hofstätter
Published: (2005)
by: Cássia Rebelo Hofstätter
Published: (2005)
Imagem e Identidade Institucional: Um Estudo Aplicado à Feevale
by: Cássia Rebello Hofstätter
Published: (2009)
by: Cássia Rebello Hofstätter
Published: (2009)
Understanding Addition and Subtraction in Transformers
by: Quirke, Philip, et al.
Published: (2024)
by: Quirke, Philip, et al.
Published: (2024)
CTRL-ALT-DECEIT: Sabotage Evaluations for Automated AI R&D
by: Ward, Francis Rhys, et al.
Published: (2025)
by: Ward, Francis Rhys, et al.
Published: (2025)
Chain-of-thought obfuscation learned from output supervision can generalise to unseen tasks
by: Hadida, Nathaniel Mitrani, et al.
Published: (2026)
by: Hadida, Nathaniel Mitrani, et al.
Published: (2026)
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
by: Veuthey, Jaime Raldua, et al.
Published: (2025)
by: Veuthey, Jaime Raldua, et al.
Published: (2025)
Rethinking CyberSecEval: An LLM-Aided Approach to Evaluation Critique
by: Hariharan, Suhas, et al.
Published: (2024)
by: Hariharan, Suhas, et al.
Published: (2024)
Lower overhead fault-tolerant building blocks for noisy quantum computers
by: Prabhu, Prithviraj
Published: (2026)
by: Prabhu, Prithviraj
Published: (2026)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
Uma contribuição da educação ambiental crítica para (des)construção do olhar sobre a seca no semiárido baiano
by: Lakshmi Juliane Vallim Hofstatter
Published: (2016)
by: Lakshmi Juliane Vallim Hofstatter
Published: (2016)
Ancient asexuality: No scandals found with novel data
by: Paulo Hofstatter, et al.
Published: (2024)
by: Paulo Hofstatter, et al.
Published: (2024)
Gravity driven traveling bore wave solutions to the free boundary incompressible Navier-Stokes equations
by: Stevenson, Noah, et al.
Published: (2025)
by: Stevenson, Noah, et al.
Published: (2025)
Hydrostatic bubbles of compressible fluid in an incompressible fluid
by: Jang, Juhi, et al.
Published: (2025)
by: Jang, Juhi, et al.
Published: (2025)
The traveling wave problem for the shallow water equations: well-posedness and the limits of vanishing viscosity and surface tension
by: Stevenson, Noah, et al.
Published: (2023)
by: Stevenson, Noah, et al.
Published: (2023)
Stationary wave solutions to two dimensional viscous shallow water equations: theory of small and large solutions
by: Stevenson, Noah, et al.
Published: (2025)
by: Stevenson, Noah, et al.
Published: (2025)
Stress Testing Deliberative Alignment for Anti-Scheming Training
by: Schoen, Bronson, et al.
Published: (2025)
by: Schoen, Bronson, et al.
Published: (2025)
Similar Items
-
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
by: van der Weij, Teun, et al.
Published: (2024) -
The Elicitation Game: Evaluating Capability Elicitation Techniques
by: Hofstätter, Felix, et al.
Published: (2025) -
Noise Injection Systemically Degrades Large Language Model Safety Guardrails
by: Shahani, Prithviraj Singh, et al.
Published: (2025) -
Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
by: Haimes, Jacob, et al.
Published: (2024) -
Extending Activation Steering to Broad Skills and Multiple Behaviours
by: van der Weij, Teun, et al.
Published: (2024)