Realistic honeypot evaluations for scheming propensity
Fuente:
arXiv
Salvato in:
| Autori principali: | Krakovna, Victoria, Lindner, David, Ho, Lewis, Farquhar, Sebastian, Shah, Rohin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Gram: Assessing sabotage propensities via automated alignment auditing
di: Lindner, David, et al.
Pubblicazione: (2026)
di: Lindner, David, et al.
Pubblicazione: (2026)
Evaluating Frontier Models for Stealth and Situational Awareness
di: Phuong, Mary, et al.
Pubblicazione: (2025)
di: Phuong, Mary, et al.
Pubblicazione: (2025)
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
di: Farquhar, Sebastian, et al.
Pubblicazione: (2025)
di: Farquhar, Sebastian, et al.
Pubblicazione: (2025)
Aligned, Orthogonal or In-conflict: When can we safely optimize Chain-of-Thought?
di: Kaufmann, Max, et al.
Pubblicazione: (2026)
di: Kaufmann, Max, et al.
Pubblicazione: (2026)
Evaluating Frontier Models for Dangerous Capabilities
di: Phuong, Mary, et al.
Pubblicazione: (2024)
di: Phuong, Mary, et al.
Pubblicazione: (2024)
A Pragmatic Way to Measure Chain-of-Thought Monitorability
di: Emmons, Scott, et al.
Pubblicazione: (2025)
di: Emmons, Scott, et al.
Pubblicazione: (2025)
An Approach to Technical AGI Safety and Security
di: Shah, Rohin, et al.
Pubblicazione: (2025)
di: Shah, Rohin, et al.
Pubblicazione: (2025)
AtP*: An efficient and scalable method for localizing LLM behaviour to components
di: Kramár, János, et al.
Pubblicazione: (2024)
di: Kramár, János, et al.
Pubblicazione: (2024)
Improving Water Quality Time-Series Prediction in Hong Kong using Sentinel-2 MSI Data and Google Earth Engine Cloud Computing
di: Sood, Rohin, et al.
Pubblicazione: (2024)
di: Sood, Rohin, et al.
Pubblicazione: (2024)
Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs
di: Easley, Eric, et al.
Pubblicazione: (2026)
di: Easley, Eric, et al.
Pubblicazione: (2026)
Consistency Training Helps Stop Sycophancy and Jailbreaks
di: Irpan, Alex, et al.
Pubblicazione: (2025)
di: Irpan, Alex, et al.
Pubblicazione: (2025)
Do Multilingual LLMs Think In English?
di: Schut, Lisa, et al.
Pubblicazione: (2025)
di: Schut, Lisa, et al.
Pubblicazione: (2025)
Improving Dictionary Learning with Gated Sparse Autoencoders
di: Rajamanoharan, Senthooran, et al.
Pubblicazione: (2024)
di: Rajamanoharan, Senthooran, et al.
Pubblicazione: (2024)
On scalable oversight with weak LLMs judging strong LLMs
di: Kenton, Zachary, et al.
Pubblicazione: (2024)
di: Kenton, Zachary, et al.
Pubblicazione: (2024)
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
di: Manvi, Rohin, et al.
Pubblicazione: (2024)
di: Manvi, Rohin, et al.
Pubblicazione: (2024)
Quantifying the Necessity of Chain of Thought through Opaque Serial Depth
di: Brown-Cohen, Jonah, et al.
Pubblicazione: (2026)
di: Brown-Cohen, Jonah, et al.
Pubblicazione: (2026)
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
di: Lieberum, Tom, et al.
Pubblicazione: (2024)
di: Lieberum, Tom, et al.
Pubblicazione: (2024)
Building Production-Ready Probes For Gemini
di: Kramár, János, et al.
Pubblicazione: (2026)
di: Kramár, János, et al.
Pubblicazione: (2026)
Evaluating the Goal-Directedness of Large Language Models
di: Everitt, Tom, et al.
Pubblicazione: (2025)
di: Everitt, Tom, et al.
Pubblicazione: (2025)
GeoLLM: Extracting Geospatial Knowledge from Large Language Models
di: Manvi, Rohin, et al.
Pubblicazione: (2023)
di: Manvi, Rohin, et al.
Pubblicazione: (2023)
Learning Safety Constraints from Demonstrations with Unknown Rewards
di: Lindner, David, et al.
Pubblicazione: (2023)
di: Lindner, David, et al.
Pubblicazione: (2023)
Large Language Models are Geographically Biased
di: Manvi, Rohin, et al.
Pubblicazione: (2024)
di: Manvi, Rohin, et al.
Pubblicazione: (2024)
Frontier Models Can Take Actions at Low Probabilities
di: Serrano, Alex, et al.
Pubblicazione: (2026)
di: Serrano, Alex, et al.
Pubblicazione: (2026)
Agent Skills Enable a New Class of Realistic and Trivially Simple Prompt Injections
di: Schmotz, David, et al.
Pubblicazione: (2025)
di: Schmotz, David, et al.
Pubblicazione: (2025)
Stress-Testing Alignment Audits With Prompt-Level Strategic Deception
di: Daniels, Oliver, et al.
Pubblicazione: (2026)
di: Daniels, Oliver, et al.
Pubblicazione: (2026)
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
di: Korbak, Tomek, et al.
Pubblicazione: (2025)
di: Korbak, Tomek, et al.
Pubblicazione: (2025)
Realistic Evaluation of Test-Time Adaptation Algorithms: Unsupervised Hyperparameter Selection
di: Cygert, Sebastian, et al.
Pubblicazione: (2024)
di: Cygert, Sebastian, et al.
Pubblicazione: (2024)
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
di: Chanin, David, et al.
Pubblicazione: (2026)
di: Chanin, David, et al.
Pubblicazione: (2026)
Predicting Fault-Ride-Through Probability of Inverter-Dominated Power Grids using Machine Learning
di: Nauck, Christian, et al.
Pubblicazione: (2024)
di: Nauck, Christian, et al.
Pubblicazione: (2024)
MISR: Measuring Instrumental Self-Reasoning in Frontier Models
di: Fronsdal, Kai, et al.
Pubblicazione: (2024)
di: Fronsdal, Kai, et al.
Pubblicazione: (2024)
Synthesizing Realistic Test Data without Breaking Privacy
di: Plein, Laura, et al.
Pubblicazione: (2026)
di: Plein, Laura, et al.
Pubblicazione: (2026)
Realistic Evaluation of Deep Partial-Label Learning Algorithms
di: Wang, Wei, et al.
Pubblicazione: (2025)
di: Wang, Wei, et al.
Pubblicazione: (2025)
Adversarial Causal Tuning for Realistic Time-series Generation
di: Gkorgkolis, Nikolaos, et al.
Pubblicazione: (2025)
di: Gkorgkolis, Nikolaos, et al.
Pubblicazione: (2025)
Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky
di: Hathidara, Ashutosh, et al.
Pubblicazione: (2025)
di: Hathidara, Ashutosh, et al.
Pubblicazione: (2025)
Revisiting the LiRA Membership Inference Attack Under Realistic Assumptions
di: Jebreel, Najeeb, et al.
Pubblicazione: (2026)
di: Jebreel, Najeeb, et al.
Pubblicazione: (2026)
The recursive scheme of clustering
di: Miniak-Górecka, Alicja, et al.
Pubblicazione: (2024)
di: Miniak-Górecka, Alicja, et al.
Pubblicazione: (2024)
SportsNGEN: Sustained Generation of Realistic Multi-player Sports Gameplay
di: Thorpe, Lachlan, et al.
Pubblicazione: (2024)
di: Thorpe, Lachlan, et al.
Pubblicazione: (2024)
Calibrating Generative AI to Produce Realistic Essays for Data Augmentation
di: Wolfe, Edward W., et al.
Pubblicazione: (2026)
di: Wolfe, Edward W., et al.
Pubblicazione: (2026)
Towards Realistic Class-Incremental Learning with Free-Flow Increments
di: Xu, Zhiming, et al.
Pubblicazione: (2026)
di: Xu, Zhiming, et al.
Pubblicazione: (2026)
Evaluating Language-Model Agents on Realistic Autonomous Tasks
di: Kinniment, Megan, et al.
Pubblicazione: (2023)
di: Kinniment, Megan, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Gram: Assessing sabotage propensities via automated alignment auditing
di: Lindner, David, et al.
Pubblicazione: (2026) -
Evaluating Frontier Models for Stealth and Situational Awareness
di: Phuong, Mary, et al.
Pubblicazione: (2025) -
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
di: Farquhar, Sebastian, et al.
Pubblicazione: (2025) -
Aligned, Orthogonal or In-conflict: When can we safely optimize Chain-of-Thought?
di: Kaufmann, Max, et al.
Pubblicazione: (2026) -
Evaluating Frontier Models for Dangerous Capabilities
di: Phuong, Mary, et al.
Pubblicazione: (2024)