On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Ball, Sarah, Gluch, Greg, Goldwasser, Shafi, Kreuter, Frauke, Reingold, Omer, Rothblum, Guy N. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Cryptographic Perspective on Mitigation vs. Detection in Machine Learning
by: Gluch, Greg, et al.
Published: (2025)
by: Gluch, Greg, et al.
Published: (2025)
How to Verify Any (Reasonable) Distribution Property: Computationally Sound Argument Systems for Distributions
by: Herman, Tal, et al.
Published: (2024)
by: Herman, Tal, et al.
Published: (2024)
PINE: Efficient Norm-Bound Verification for Secret-Shared Vectors
by: Rothblum, Guy N., et al.
Published: (2023)
by: Rothblum, Guy N., et al.
Published: (2023)
Deniable Encryption in a Quantum World
by: Coladangelo, Andrea, et al.
Published: (2021)
by: Coladangelo, Andrea, et al.
Published: (2021)
To share or not to share: What risks would laypeople accept to give sensitive data to differentially-private NLP systems?
by: Weiss, Christopher, et al.
Published: (2023)
by: Weiss, Christopher, et al.
Published: (2023)
Planting Undetectable Backdoors in Machine Learning Models
by: Goldwasser, Shafi, et al.
Published: (2022)
by: Goldwasser, Shafi, et al.
Published: (2022)
Oblivious Defense in ML Models: Backdoor Removal without Detection
by: Goldwasser, Shafi, et al.
Published: (2024)
by: Goldwasser, Shafi, et al.
Published: (2024)
Accuracy vs. Accuracy: Computational Tradeoffs Between Classification Rates and Utility
by: Amit, Noga, et al.
Published: (2025)
by: Amit, Noga, et al.
Published: (2025)
AI Native Asset Intelligence
by: Engelberg, Gal, et al.
Published: (2026)
by: Engelberg, Gal, et al.
Published: (2026)
Local Pan-Privacy for Federated Analytics
by: Feldman, Vitaly, et al.
Published: (2025)
by: Feldman, Vitaly, et al.
Published: (2025)
Efficient Public Verification of Private ML via Regularization
by: Bell, Zoë Ruha, et al.
Published: (2025)
by: Bell, Zoë Ruha, et al.
Published: (2025)
Blockchain and AI: Securing Intelligent Networks for the Future
by: Dutta, Joy, et al.
Published: (2026)
by: Dutta, Joy, et al.
Published: (2026)
Structural Enforcement of Goal Integrity in AI Agents via Separation-of-Powers Architecture
by: Xiang, Rong
Published: (2026)
by: Xiang, Rong
Published: (2026)
PREAMBLE: Private and Efficient Aggregation via Block Sparse Vectors
by: Asi, Hilal, et al.
Published: (2025)
by: Asi, Hilal, et al.
Published: (2025)
Privacy-Preserving Decentralized AI with Confidential Computing
by: Lee, Dayeol, et al.
Published: (2024)
by: Lee, Dayeol, et al.
Published: (2024)
Research on Enhancing Cloud Computing Network Security using Artificial Intelligence Algorithms
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Unlocking Apple's Private Cloud Compute: An Analysis of Privacy-Preserving Artificial Intelligence
by: Dittmar, Yannik, et al.
Published: (2026)
by: Dittmar, Yannik, et al.
Published: (2026)
Evasive Intelligence: Lessons from Malware Analysis for Evaluating AI Agents
by: Aonzo, Simone, et al.
Published: (2026)
by: Aonzo, Simone, et al.
Published: (2026)
ATLANTIS: AI-driven Threat Localization, Analysis, and Triage Intelligence System
by: Kim, Taesoo, et al.
Published: (2025)
by: Kim, Taesoo, et al.
Published: (2025)
Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
by: Kim, Jinhwa, et al.
Published: (2025)
by: Kim, Jinhwa, et al.
Published: (2025)
AI-Powered Algorithms for the Prevention and Detection of Computer Malware Infections
by: Keshava, Rakesh, et al.
Published: (2026)
by: Keshava, Rakesh, et al.
Published: (2026)
The Global Impact of AI-Artificial Intelligence: Recent Advances and Future Directions, A Review
by: Pachegowda, Chandregowda
Published: (2023)
by: Pachegowda, Chandregowda
Published: (2023)
Trojans in Artificial Intelligence (TrojAI) Final Report
by: Reese, Kristopher W., et al.
Published: (2026)
by: Reese, Kristopher W., et al.
Published: (2026)
Artificial Intelligence for Secured Information Systems in Smart Cities: Collaborative IoT Computing with Deep Reinforcement Learning and Blockchain
by: Far, Amin Zakaie, et al.
Published: (2024)
by: Far, Amin Zakaie, et al.
Published: (2024)
EdgeShield: A Universal and Efficient Edge Computing Framework for Robust AI
by: Zhong, Duo, et al.
Published: (2024)
by: Zhong, Duo, et al.
Published: (2024)
When Agents Handle Secrets: A Survey of Confidential Computing for Agentic AI
by: Forough, Javad, et al.
Published: (2026)
by: Forough, Javad, et al.
Published: (2026)
Cyber Threat Intelligence for Artificial Intelligence Systems
by: Krawczyk, Natalia, et al.
Published: (2026)
by: Krawczyk, Natalia, et al.
Published: (2026)
Reimagining Safety Alignment with An Image
by: Xia, Yifan, et al.
Published: (2025)
by: Xia, Yifan, et al.
Published: (2025)
AI-Driven Security in Cloud Computing: Enhancing Threat Detection, Automated Response, and Cyber Resilience
by: Shaffi, Shamnad Mohamed, et al.
Published: (2025)
by: Shaffi, Shamnad Mohamed, et al.
Published: (2025)
Evaluating the Reliability and Fidelity of Automated Judgment Systems of Large Language Models
by: Biskupski, Tom, et al.
Published: (2026)
by: Biskupski, Tom, et al.
Published: (2026)
Containment Verification: AI Safety Guarantees Independent of Alignment
by: Moon, Royce, et al.
Published: (2026)
by: Moon, Royce, et al.
Published: (2026)
CBPF: Filtering Poisoned Data Based on Composite Backdoor Attack
by: Xia, Hanfeng, et al.
Published: (2024)
by: Xia, Hanfeng, et al.
Published: (2024)
Securing Generative AI in Healthcare: A Zero-Trust Architecture Powered by Confidential Computing on Google Cloud
by: Amanna, Adaobi, et al.
Published: (2025)
by: Amanna, Adaobi, et al.
Published: (2025)
Emission Impossible: privacy-preserving carbon emissions claims
by: Man, Jessica, et al.
Published: (2025)
by: Man, Jessica, et al.
Published: (2025)
Bitcoin-Enhanced Proof-of-Stake Security: Possibilities and Impossibilities
by: Tas, Ertem Nusret, et al.
Published: (2022)
by: Tas, Ertem Nusret, et al.
Published: (2022)
AI-Driven Cyber Threat Intelligence Automation
by: Shah, Shrit, et al.
Published: (2024)
by: Shah, Shrit, et al.
Published: (2024)
What Was Your Prompt? A Remote Keylogging Attack on AI Assistants
by: Weiss, Roy, et al.
Published: (2024)
by: Weiss, Roy, et al.
Published: (2024)
Agent Safety Alignment via Reinforcement Learning
by: Sha, Zeyang, et al.
Published: (2025)
by: Sha, Zeyang, et al.
Published: (2025)
EVA: Editing for Versatile Alignment against Jailbreaks
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
UK AISI Alignment Evaluation Case-Study
by: Souly, Alexandra, et al.
Published: (2026)
by: Souly, Alexandra, et al.
Published: (2026)
Similar Items
-
A Cryptographic Perspective on Mitigation vs. Detection in Machine Learning
by: Gluch, Greg, et al.
Published: (2025) -
How to Verify Any (Reasonable) Distribution Property: Computationally Sound Argument Systems for Distributions
by: Herman, Tal, et al.
Published: (2024) -
PINE: Efficient Norm-Bound Verification for Secret-Shared Vectors
by: Rothblum, Guy N., et al.
Published: (2023) -
Deniable Encryption in a Quantum World
by: Coladangelo, Andrea, et al.
Published: (2021) -
To share or not to share: What risks would laypeople accept to give sensitive data to differentially-private NLP systems?
by: Weiss, Christopher, et al.
Published: (2023)