Password-Activated Shutdown Protocols for Misaligned Frontier Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Williams, Kai, Subramani, Rohan, Ward, Francis Rhys |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The New Frontier of Cybersecurity: Emerging Threats and Innovations
por: Dave, Daksh, et al.
Publicado: (2023)
por: Dave, Daksh, et al.
Publicado: (2023)
Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots
por: Shayegani, Erfan, et al.
Publicado: (2025)
por: Shayegani, Erfan, et al.
Publicado: (2025)
Robust Safety Monitoring of Language Models via Activation Watermarking
por: Aremu, Toluwani, et al.
Publicado: (2026)
por: Aremu, Toluwani, et al.
Publicado: (2026)
Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
por: Shayegani, Erfan, et al.
Publicado: (2025)
por: Shayegani, Erfan, et al.
Publicado: (2025)
MAYA: Addressing Inconsistencies in Generative Password Guessing through a Unified Benchmark
por: Corrias, William, et al.
Publicado: (2025)
por: Corrias, William, et al.
Publicado: (2025)
When Intelligence Fails: An Empirical Study on Why LLMs Struggle with Password Cracking
por: Rehman, Mohammad Abdul, et al.
Publicado: (2025)
por: Rehman, Mohammad Abdul, et al.
Publicado: (2025)
Persona-Model Collapse in Emergent Misalignment
por: Costa, Davi Bastos, et al.
Publicado: (2026)
por: Costa, Davi Bastos, et al.
Publicado: (2026)
Confidential Guardian: Cryptographically Prohibiting the Abuse of Model Abstention
por: Rabanser, Stephan, et al.
Publicado: (2025)
por: Rabanser, Stephan, et al.
Publicado: (2025)
Adversarial Augmentation and Active Sampling for Robust Cyber Anomaly Detection
por: Benabderrahmane, Sidahmed, et al.
Publicado: (2025)
por: Benabderrahmane, Sidahmed, et al.
Publicado: (2025)
Robustness and Cybersecurity in the EU Artificial Intelligence Act
por: Nolte, Henrik, et al.
Publicado: (2025)
por: Nolte, Henrik, et al.
Publicado: (2025)
FAIRPLAI: A Human-in-the-Loop Approach to Fair and Private Machine Learning
por: Sanchez Jr., David, et al.
Publicado: (2025)
por: Sanchez Jr., David, et al.
Publicado: (2025)
Secure Multi-Modal Data Fusion in Federated Digital Health Systems via MCP
por: Aueawatthanaphisut, Aueaphum
Publicado: (2025)
por: Aueawatthanaphisut, Aueaphum
Publicado: (2025)
VISION: Robust and Interpretable Code Vulnerability Detection Leveraging Counterfactual Augmentation
por: Egea, David, et al.
Publicado: (2025)
por: Egea, David, et al.
Publicado: (2025)
Unifying Re-Identification, Attribute Inference, and Data Reconstruction Risks in Differential Privacy
por: Kulynych, Bogdan, et al.
Publicado: (2025)
por: Kulynych, Bogdan, et al.
Publicado: (2025)
The Wolf Within: Covert Injection of Malice into MLLM Societies via an MLLM Operative
por: Tan, Zhen, et al.
Publicado: (2024)
por: Tan, Zhen, et al.
Publicado: (2024)
Making AI-Assisted Grant Evaluation Auditable without Exposing the Model
por: Bicakci, Kemal
Publicado: (2026)
por: Bicakci, Kemal
Publicado: (2026)
Inferring Discussion Topics about Exploitation of Vulnerabilities from Underground Hacking Forums
por: Moreno-Vera, Felipe
Publicado: (2024)
por: Moreno-Vera, Felipe
Publicado: (2024)
Trustless Audits without Revealing Data or Models
por: Waiwitlikhit, Suppakit, et al.
Publicado: (2024)
por: Waiwitlikhit, Suppakit, et al.
Publicado: (2024)
Decentralized autonomous organization and blockchain-based incentivization framework for community-based facilities management
por: Ly, Reachsak, et al.
Publicado: (2026)
por: Ly, Reachsak, et al.
Publicado: (2026)
Watermarking Should Be Treated as a Monitoring Primitive
por: Aremu, Toluwani, et al.
Publicado: (2026)
por: Aremu, Toluwani, et al.
Publicado: (2026)
A Survey of Privacy-Preserving Model Explanations: Privacy Risks, Attacks, and Countermeasures
por: Nguyen, Thanh Tam, et al.
Publicado: (2024)
por: Nguyen, Thanh Tam, et al.
Publicado: (2024)
Privacy-hardened and hallucination-resistant synthetic data generation with logic-solvers
por: Burgess, Mark A., et al.
Publicado: (2024)
por: Burgess, Mark A., et al.
Publicado: (2024)
NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
por: Shao, Minghao, et al.
Publicado: (2024)
por: Shao, Minghao, et al.
Publicado: (2024)
SoK: On the Offensive Potential of AI
por: Schröer, Saskia Laura, et al.
Publicado: (2024)
por: Schröer, Saskia Laura, et al.
Publicado: (2024)
Benchmark Early and Red Team Often: A Framework for Assessing and Managing Dual-Use Hazards of AI Foundation Models
por: Barrett, Anthony M., et al.
Publicado: (2024)
por: Barrett, Anthony M., et al.
Publicado: (2024)
SecGenAI: Enhancing Security of Cloud-based Generative AI Applications within Australian Critical Technologies of National Interest
por: Haryanto, Christoforus Yoga, et al.
Publicado: (2024)
por: Haryanto, Christoforus Yoga, et al.
Publicado: (2024)
Hidden Poison: Machine Unlearning Enables Camouflaged Poisoning Attacks
por: Di, Jimmy Z., et al.
Publicado: (2022)
por: Di, Jimmy Z., et al.
Publicado: (2022)
Differentially Private Data Release on Graphs: Inefficiencies and Unfairness
por: Fioretto, Ferdinando, et al.
Publicado: (2024)
por: Fioretto, Ferdinando, et al.
Publicado: (2024)
Privacy Bias in Language Models: A Contextual Integrity-based Auditing Metric
por: Shvartzshnaider, Yan, et al.
Publicado: (2024)
por: Shvartzshnaider, Yan, et al.
Publicado: (2024)
A Public Theory of Distillation Resistance via Constraint-Coupled Reasoning Architectures
por: Wei, Peng, et al.
Publicado: (2026)
por: Wei, Peng, et al.
Publicado: (2026)
Machine Unlearning Fails to Remove Data Poisoning Attacks
por: Pawelczyk, Martin, et al.
Publicado: (2024)
por: Pawelczyk, Martin, et al.
Publicado: (2024)
Knowledge Distillation-Based Model Extraction Attack using GAN-based Private Counterfactual Explanations
por: Ezzeddine, Fatima, et al.
Publicado: (2024)
por: Ezzeddine, Fatima, et al.
Publicado: (2024)
Privacy at a Price: Exploring its Dual Impact on AI Fairness
por: Yang, Mengmeng, et al.
Publicado: (2024)
por: Yang, Mengmeng, et al.
Publicado: (2024)
PUFFLE: Balancing Privacy, Utility, and Fairness in Federated Learning
por: Corbucci, Luca, et al.
Publicado: (2024)
por: Corbucci, Luca, et al.
Publicado: (2024)
Personal Information Parroting in Language Models
por: Subramani, Nishant, et al.
Publicado: (2026)
por: Subramani, Nishant, et al.
Publicado: (2026)
Frontier AI's Impact on the Cybersecurity Landscape
por: Potter, Yujin, et al.
Publicado: (2025)
por: Potter, Yujin, et al.
Publicado: (2025)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
por: Wang, Yifei, et al.
Publicado: (2024)
por: Wang, Yifei, et al.
Publicado: (2024)
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
por: Betley, Jan, et al.
Publicado: (2025)
por: Betley, Jan, et al.
Publicado: (2025)
Agentic Misalignment: How LLMs Could Be Insider Threats
por: Lynch, Aengus, et al.
Publicado: (2025)
por: Lynch, Aengus, et al.
Publicado: (2025)
Clear, Compelling Arguments: Rethinking the Foundations of Frontier AI Safety Cases
por: Feakins, Shaun, et al.
Publicado: (2026)
por: Feakins, Shaun, et al.
Publicado: (2026)
Ejemplares similares
-
The New Frontier of Cybersecurity: Emerging Threats and Innovations
por: Dave, Daksh, et al.
Publicado: (2023) -
Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots
por: Shayegani, Erfan, et al.
Publicado: (2025) -
Robust Safety Monitoring of Language Models via Activation Watermarking
por: Aremu, Toluwani, et al.
Publicado: (2026) -
Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
por: Shayegani, Erfan, et al.
Publicado: (2025) -
MAYA: Addressing Inconsistencies in Generative Password Guessing through a Unified Benchmark
por: Corrias, William, et al.
Publicado: (2025)