Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wen, Jiaxin, Hebbar, Vivek, Larson, Caleb, Bhatt, Aryan, Radhakrishnan, Ansh, Sharma, Mrinank, Sleight, Henry, Feng, Shi, He, He, Perez, Ethan, Shlegeris, Buck, Khan, Akbir |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases
par: Gan, Eric, et autres
Publié: (2026)
par: Gan, Eric, et autres
Publié: (2026)
Evaluating Control Protocols for Untrusted AI Agents
par: Kutasov, Jon, et autres
Publié: (2025)
par: Kutasov, Jon, et autres
Publié: (2025)
Ctrl-Z: Controlling AI Agents via Resampling
par: Bhatt, Aryan, et autres
Publié: (2025)
par: Bhatt, Aryan, et autres
Publié: (2025)
Rapid Response: Mitigating LLM Jailbreaks with a Few Examples
par: Peng, Alwin, et autres
Publié: (2024)
par: Peng, Alwin, et autres
Publié: (2024)
Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks
par: Youstra, Jack, et autres
Publié: (2025)
par: Youstra, Jack, et autres
Publié: (2025)
Debating with More Persuasive LLMs Leads to More Truthful Answers
par: Khan, Akbir, et autres
Publié: (2024)
par: Khan, Akbir, et autres
Publié: (2024)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
par: Griffin, Charlie, et autres
Publié: (2024)
par: Griffin, Charlie, et autres
Publié: (2024)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
par: Sheshadri, Abhay, et autres
Publié: (2024)
par: Sheshadri, Abhay, et autres
Publié: (2024)
Agentic Misalignment: How LLMs Could Be Insider Threats
par: Lynch, Aengus, et autres
Publié: (2025)
par: Lynch, Aengus, et autres
Publié: (2025)
Best-of-N Jailbreaking
par: Hughes, John, et autres
Publié: (2024)
par: Hughes, John, et autres
Publié: (2024)
Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
par: Wang, Tony T., et autres
Publié: (2024)
par: Wang, Tony T., et autres
Publié: (2024)
Language Models Learn to Mislead Humans via RLHF
par: Wen, Jiaxin, et autres
Publié: (2024)
par: Wen, Jiaxin, et autres
Publié: (2024)
AI Control: Improving Safety Despite Intentional Subversion
par: Greenblatt, Ryan, et autres
Publié: (2023)
par: Greenblatt, Ryan, et autres
Publié: (2023)
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
par: Korbak, Tomek, et autres
Publié: (2025)
par: Korbak, Tomek, et autres
Publié: (2025)
Mock Theta Functions as Optimal Stopping Criteria for Photonic Quantum Entropy Computation
par: Ansh Sharma, Ansh, et autres
Publié: (2026)
par: Ansh Sharma, Ansh, et autres
Publié: (2026)
Removing Sandbagging in LLMs by Training with Weak Supervision
par: Ryd, Emil, et autres
Publié: (2026)
par: Ryd, Emil, et autres
Publié: (2026)
Factorio Learning Environment
par: Hopkins, Jack, et autres
Publié: (2025)
par: Hopkins, Jack, et autres
Publié: (2025)
Language models are better than humans at next-token prediction
par: Shlegeris, Buck, et autres
Publié: (2022)
par: Shlegeris, Buck, et autres
Publié: (2022)
Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?
par: Mallen, Alex, et autres
Publié: (2024)
par: Mallen, Alex, et autres
Publié: (2024)
Polysemanticity and Capacity in Neural Networks
par: Scherlis, Adam, et autres
Publié: (2022)
par: Scherlis, Adam, et autres
Publié: (2022)
A sketch of an AI control safety case
par: Korbak, Tomek, et autres
Publié: (2025)
par: Korbak, Tomek, et autres
Publié: (2025)
Characterizing Paraphrase-Induced Failures in Lean 4 Autoformalization
par: Feng, William, et autres
Publié: (2026)
par: Feng, William, et autres
Publié: (2026)
The Duty of Knowing Oneself as One Appears: A Response to Kant’s Problem of Moral Self-Knowledge
par: Vivek Kumar Radhakrishnan
Publié: (2019)
par: Vivek Kumar Radhakrishnan
Publié: (2019)
LLMs as Debate Partners: Utilizing Genetic Algorithms and Adversarial Search for Adaptive Arguments
par: Aryan, Prakash
Publié: (2024)
par: Aryan, Prakash
Publié: (2024)
Beware Untrusted Simulators -- Reward-Free Backdoor Attacks in Reinforcement Learning
par: Rathbun, Ethan, et autres
Publié: (2026)
par: Rathbun, Ethan, et autres
Publié: (2026)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
par: Kutasov, Jonathan, et autres
Publié: (2025)
par: Kutasov, Jonathan, et autres
Publié: (2025)
Failures to Find Transferable Image Jailbreaks Between Vision-Language Models
par: Schaeffer, Rylan, et autres
Publié: (2024)
par: Schaeffer, Rylan, et autres
Publié: (2024)
All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
par: Guo, Shiyuan, et autres
Publié: (2025)
par: Guo, Shiyuan, et autres
Publié: (2025)
The LLM Has Left The Chat: Evidence of Bail Preferences in Large Language Models
par: Ensign, Danielle, et autres
Publié: (2025)
par: Ensign, Danielle, et autres
Publié: (2025)
Programming by Backprop: An Instruction is Worth 100 Examples When Finetuning LLMs
par: Cook, Jonathan, et autres
Publié: (2025)
par: Cook, Jonathan, et autres
Publié: (2025)
The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?
par: Hägele, Alexander, et autres
Publié: (2026)
par: Hägele, Alexander, et autres
Publié: (2026)
PragLocker: Protecting Agent Intellectual Property in Untrusted Deployments via Non-Portable Prompts
par: Li, Qinfeng, et autres
Publié: (2026)
par: Li, Qinfeng, et autres
Publié: (2026)
Unsupervised Elicitation of Language Models
par: Wen, Jiaxin, et autres
Publié: (2025)
par: Wen, Jiaxin, et autres
Publié: (2025)
UntrustVul: Improving the Usability of Vulnerability Detection Models by Reducing Untrustworthy Alerts
par: Anonymous, Anonymous
Publié: (2025)
par: Anonymous, Anonymous
Publié: (2025)
Hybrid Implementation for Untrusted-node-based Quantum Key Distribution Network
par: Liu, Jingyang, et autres
Publié: (2025)
par: Liu, Jingyang, et autres
Publié: (2025)
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
par: Hubinger, Evan, et autres
Publié: (2024)
par: Hubinger, Evan, et autres
Publié: (2024)
Who's in Charge? Disempowerment Patterns in Real-World LLM Usage
par: Sharma, Mrinank, et autres
Publié: (2026)
par: Sharma, Mrinank, et autres
Publié: (2026)
Incorporating Unlabelled Data into Bayesian Neural Networks
par: Sharma, Mrinank, et autres
Publié: (2023)
par: Sharma, Mrinank, et autres
Publié: (2023)
$κ$-solutions with the round cylinder as an asymptotic shrinker
par: Hebbar, Aprameya Girish
Publié: (2026)
par: Hebbar, Aprameya Girish
Publié: (2026)
Moving Faster and Reducing Risk: Using LLMs in Release Deployment
par: Abreu, Rui, et autres
Publié: (2024)
par: Abreu, Rui, et autres
Publié: (2024)
Documents similaires
-
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases
par: Gan, Eric, et autres
Publié: (2026) -
Evaluating Control Protocols for Untrusted AI Agents
par: Kutasov, Jon, et autres
Publié: (2025) -
Ctrl-Z: Controlling AI Agents via Resampling
par: Bhatt, Aryan, et autres
Publié: (2025) -
Rapid Response: Mitigating LLM Jailbreaks with a Few Examples
par: Peng, Alwin, et autres
Publié: (2024) -
Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks
par: Youstra, Jack, et autres
Publié: (2025)