Intent Laundering: AI Safety Datasets Are Not What They Seem
Fuente:
arXiv
Saved in:
| Main Authors: | Golchin, Shahriar, Wetter, Marc |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Time Travel in LLMs: Tracing Data Contamination in Large Language Models
by: Golchin, Shahriar, et al.
Published: (2023)
by: Golchin, Shahriar, et al.
Published: (2023)
What is in Your Safe Data? Identifying Benign Data that Breaks Safety
by: He, Luxi, et al.
Published: (2024)
by: He, Luxi, et al.
Published: (2024)
An Adversarial Perspective on Machine Unlearning for AI Safety
by: Łucki, Jakub, et al.
Published: (2024)
by: Łucki, Jakub, et al.
Published: (2024)
Trustworthy AI: Safety, Bias, and Privacy -- A Survey
by: Fang, Xingli, et al.
Published: (2025)
by: Fang, Xingli, et al.
Published: (2025)
Adversarial Intent is a Latent Variable: Stateful Trust Inference for Securing Multimodal Agentic RAG
by: Singh, Inderjeet, et al.
Published: (2026)
by: Singh, Inderjeet, et al.
Published: (2026)
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies
by: Parmar, Manojkumar, et al.
Published: (2025)
by: Parmar, Manojkumar, et al.
Published: (2025)
Safety Alignment Can Be Not Superficial With Explicit Safety Signals
by: Li, Jianwei, et al.
Published: (2025)
by: Li, Jianwei, et al.
Published: (2025)
AnnoCTR: A Dataset for Detecting and Linking Entities, Tactics, and Techniques in Cyber Threat Reports
by: Lange, Lukas, et al.
Published: (2024)
by: Lange, Lukas, et al.
Published: (2024)
Lifelong Safety Alignment for Language Models
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Certifying LLM Safety against Adversarial Prompting
by: Kumar, Aounon, et al.
Published: (2023)
by: Kumar, Aounon, et al.
Published: (2023)
On the Role of Attention Heads in Large Language Model Safety
by: Zhou, Zhenhong, et al.
Published: (2024)
by: Zhou, Zhenhong, et al.
Published: (2024)
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
by: Vega, Jason, et al.
Published: (2023)
by: Vega, Jason, et al.
Published: (2023)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
by: Wu, Tong, et al.
Published: (2024)
by: Wu, Tong, et al.
Published: (2024)
Probing the Robustness of Large Language Models Safety to Latent Perturbations
by: Gu, Tianle, et al.
Published: (2025)
by: Gu, Tianle, et al.
Published: (2025)
MANATEE: Inference-Time Lightweight Diffusion Based Safety Defense for LLMs
by: Kan, Chun Yan Ryan, et al.
Published: (2026)
by: Kan, Chun Yan Ryan, et al.
Published: (2026)
Amplification Effects in Test-Time Reinforcement Learning: Safety and Reasoning Vulnerabilities
by: Khattar, Vanshaj, et al.
Published: (2026)
by: Khattar, Vanshaj, et al.
Published: (2026)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
by: Chen, Taiye, et al.
Published: (2025)
by: Chen, Taiye, et al.
Published: (2025)
In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement
by: Shetty, Anudeex, et al.
Published: (2026)
by: Shetty, Anudeex, et al.
Published: (2026)
SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
by: Li, Lijun, et al.
Published: (2024)
by: Li, Lijun, et al.
Published: (2024)
Exposing LLM Safety Gaps Through Mathematical Encoding:New Attacks and Systematic Analysis
by: Zhang, Haoyu, et al.
Published: (2026)
by: Zhang, Haoyu, et al.
Published: (2026)
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
by: Wang, Kun, et al.
Published: (2025)
by: Wang, Kun, et al.
Published: (2025)
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs
by: Fonseca, Joao, et al.
Published: (2025)
by: Fonseca, Joao, et al.
Published: (2025)
EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors
by: Banayeeanzade, Amin, et al.
Published: (2026)
by: Banayeeanzade, Amin, et al.
Published: (2026)
IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
by: Guo, Chuan, et al.
Published: (2026)
by: Guo, Chuan, et al.
Published: (2026)
ObfuscaTune: Obfuscated Offsite Fine-tuning and Inference of Proprietary LLMs on Private Datasets
by: Frikha, Ahmed, et al.
Published: (2024)
by: Frikha, Ahmed, et al.
Published: (2024)
SecureCode: A Production-Grade Multi-Turn Dataset for Training Security-Aware Code Generation Models
by: Thornton, Scott
Published: (2025)
by: Thornton, Scott
Published: (2025)
Cross-Session Threats in AI Agents: Benchmark, Evaluation, and Algorithms
by: Azarafrooz, Ari
Published: (2026)
by: Azarafrooz, Ari
Published: (2026)
UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
by: Shumailov, Ilia, et al.
Published: (2024)
by: Shumailov, Ilia, et al.
Published: (2024)
Superficial Safety Alignment Hypothesis
by: Li, Jianwei, et al.
Published: (2024)
by: Li, Jianwei, et al.
Published: (2024)
Evading Data Contamination Detection for Language Models is (too) Easy
by: Dekoninck, Jasper, et al.
Published: (2024)
by: Dekoninck, Jasper, et al.
Published: (2024)
A Large-Scale Empirical Analysis of Custom GPTs' Vulnerabilities in the OpenAI Ecosystem
by: Ogundoyin, Sunday Oyinlola, et al.
Published: (2025)
by: Ogundoyin, Sunday Oyinlola, et al.
Published: (2025)
BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
by: Zhang, Andy K., et al.
Published: (2025)
by: Zhang, Andy K., et al.
Published: (2025)
Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
by: Labiad, Ismail, et al.
Published: (2025)
by: Labiad, Ismail, et al.
Published: (2025)
When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?
by: Tian, Yuan, et al.
Published: (2026)
by: Tian, Yuan, et al.
Published: (2026)
What Matters For Safety Alignment?
by: Li, Xing, et al.
Published: (2026)
by: Li, Xing, et al.
Published: (2026)
Deep Learning Approaches for Anti-Money Laundering on Mobile Transactions: Review, Framework, and Directions
by: Fan, Jiani, et al.
Published: (2025)
by: Fan, Jiani, et al.
Published: (2025)
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
by: Jin, Chang, et al.
Published: (2026)
by: Jin, Chang, et al.
Published: (2026)
What Makes an Evaluation Useful? Common Pitfalls and Best Practices
by: Gekker, Gil, et al.
Published: (2025)
by: Gekker, Gil, et al.
Published: (2025)
Generative AI Security: Challenges and Countermeasures
by: Zhu, Banghua, et al.
Published: (2024)
by: Zhu, Banghua, et al.
Published: (2024)
Urania: Differentially Private Insights into AI Use
by: Liu, Daogao, et al.
Published: (2025)
by: Liu, Daogao, et al.
Published: (2025)
Similar Items
-
Time Travel in LLMs: Tracing Data Contamination in Large Language Models
by: Golchin, Shahriar, et al.
Published: (2023) -
What is in Your Safe Data? Identifying Benign Data that Breaks Safety
by: He, Luxi, et al.
Published: (2024) -
An Adversarial Perspective on Machine Unlearning for AI Safety
by: Łucki, Jakub, et al.
Published: (2024) -
Trustworthy AI: Safety, Bias, and Privacy -- A Survey
by: Fang, Xingli, et al.
Published: (2025) -
Adversarial Intent is a Latent Variable: Stateful Trust Inference for Securing Multimodal Agentic RAG
by: Singh, Inderjeet, et al.
Published: (2026)