AuthorMist: Evading AI Text Detectors with Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | David, Isaac, Gervais, Arthur |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evading Deep Learning-Based Malware Detectors via Obfuscation: A Deep Reinforcement Learning Approach
by: Etter, Brian, et al.
Published: (2024)
by: Etter, Brian, et al.
Published: (2024)
Multi-Agent Penetration Testing AI for the Web
by: David, Isaac, et al.
Published: (2025)
by: David, Isaac, et al.
Published: (2025)
StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors
by: Ranganath, Suraj, et al.
Published: (2026)
by: Ranganath, Suraj, et al.
Published: (2026)
Patch2Vuln: Agentic Reconstruction of Vulnerabilities from Linux Distribution Binary Patches
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
Measuring Safety Alignment Effects in Autonomous Security Agents
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
TH-Bench: Evaluating Evading Attacks via Humanizing AI Text on Machine-Generated Text Detectors
by: Zheng, Jingyi, et al.
Published: (2025)
by: Zheng, Jingyi, et al.
Published: (2025)
AI Agent Smart Contract Exploit Generation
by: Gervais, Arthur, et al.
Published: (2025)
by: Gervais, Arthur, et al.
Published: (2025)
Learning to Watermark LLM-generated Text via Reinforcement Learning
by: Xu, Xiaojun, et al.
Published: (2024)
by: Xu, Xiaojun, et al.
Published: (2024)
Evading Data Contamination Detection for Language Models is (too) Easy
by: Dekoninck, Jasper, et al.
Published: (2024)
by: Dekoninck, Jasper, et al.
Published: (2024)
Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment
by: Zhang, Peng, et al.
Published: (2025)
by: Zhang, Peng, et al.
Published: (2025)
RoBCtrl: Attacking GNN-Based Social Bot Detectors via Reinforced Manipulation of Bots Control Interaction
by: Yang, Yingguang, et al.
Published: (2025)
by: Yang, Yingguang, et al.
Published: (2025)
Towards Optimal Agentic Architectures for Offensive Security Tasks
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
BAFFLE: Hiding Backdoors in Offline Reinforcement Learning Datasets
by: Gong, Chen, et al.
Published: (2022)
by: Gong, Chen, et al.
Published: (2022)
Adversarial Attacks on Transformers-Based Malware Detectors
by: Jakhotiya, Yash, et al.
Published: (2022)
by: Jakhotiya, Yash, et al.
Published: (2022)
Optimal Zero-Shot Detector for Multi-Armed Attacks
by: Granese, Federica, et al.
Published: (2024)
by: Granese, Federica, et al.
Published: (2024)
Syntax- and Compilation-Preserving Evasion of LLM Vulnerability Detectors
by: Sun, Luze, et al.
Published: (2026)
by: Sun, Luze, et al.
Published: (2026)
Concept Drift Adaptation Using Self-Supervised and Reinforcement Learning In Android Malware Detection
by: Sabbah, Ahmed, et al.
Published: (2026)
by: Sabbah, Ahmed, et al.
Published: (2026)
Autonomous Network Defence using Reinforcement Learning
by: Foley, Myles, et al.
Published: (2024)
by: Foley, Myles, et al.
Published: (2024)
TBDetector:Transformer-Based Detector for Advanced Persistent Threats with Provenance Graph
by: Wang, Nan, et al.
Published: (2023)
by: Wang, Nan, et al.
Published: (2023)
AgentGuardian: Learning Access Control Policies to Govern AI Agent Behavior
by: Abaev, Nadya, et al.
Published: (2026)
by: Abaev, Nadya, et al.
Published: (2026)
Differentially Private Deep Model-Based Reinforcement Learning
by: Rio, Alexandre, et al.
Published: (2024)
by: Rio, Alexandre, et al.
Published: (2024)
Near-Optimal Reinforcement Learning with Shuffle Differential Privacy
by: Bai, Shaojie, et al.
Published: (2024)
by: Bai, Shaojie, et al.
Published: (2024)
RICE: Breaking Through the Training Bottlenecks of Reinforcement Learning with Explanation
by: Cheng, Zelei, et al.
Published: (2024)
by: Cheng, Zelei, et al.
Published: (2024)
Mitigating Deep Reinforcement Learning Backdoors in the Neural Activation Space
by: Vyas, Sanyam, et al.
Published: (2024)
by: Vyas, Sanyam, et al.
Published: (2024)
A General Black-box Adversarial Attack on Graph-based Fake News Detectors
by: Zhu, Peican, et al.
Published: (2024)
by: Zhu, Peican, et al.
Published: (2024)
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
Modeling Behavioral Preferences of Cyber Adversaries Using Inverse Reinforcement Learning
by: Shinde, Aditya, et al.
Published: (2025)
by: Shinde, Aditya, et al.
Published: (2025)
Deep Reinforcement Learning for Phishing Detection with Transformer-Based Semantic Features
by: Faisal, Aseer Al
Published: (2025)
by: Faisal, Aseer Al
Published: (2025)
FRAMU: Attention-based Machine Unlearning using Federated Reinforcement Learning
by: Shaik, Thanveer, et al.
Published: (2023)
by: Shaik, Thanveer, et al.
Published: (2023)
Advanced Persistent Threats (APT) Attribution Using Deep Reinforcement Learning
by: Basnet, Animesh Singh, et al.
Published: (2024)
by: Basnet, Animesh Singh, et al.
Published: (2024)
Optimizing Cyber Defense in Dynamic Active Directories through Reinforcement Learning
by: Goel, Diksha, et al.
Published: (2024)
by: Goel, Diksha, et al.
Published: (2024)
Leveraging Reinforcement Learning in Red Teaming for Advanced Ransomware Attack Simulations
by: Wang, Cheng, et al.
Published: (2024)
by: Wang, Cheng, et al.
Published: (2024)
Knowledge-Informed Auto-Penetration Testing Based on Reinforcement Learning with Reward Machine
by: Li, Yuanliang, et al.
Published: (2024)
by: Li, Yuanliang, et al.
Published: (2024)
Universal Black-Box Reward Poisoning Attack against Offline Reinforcement Learning
by: Xu, Yinglun, et al.
Published: (2024)
by: Xu, Yinglun, et al.
Published: (2024)
PoolFlip: A Multi-Agent Reinforcement Learning Security Environment for Cyber Defense
by: Cadet, Xavier, et al.
Published: (2025)
by: Cadet, Xavier, et al.
Published: (2025)
UNIDOOR: A Universal Framework for Action-Level Backdoor Attacks in Deep Reinforcement Learning
by: Ma, Oubo, et al.
Published: (2025)
by: Ma, Oubo, et al.
Published: (2025)
Enhancing Security in Deep Reinforcement Learning: A Comprehensive Survey on Adversarial Attacks and Defenses
by: Yichao, Wu, et al.
Published: (2025)
by: Yichao, Wu, et al.
Published: (2025)
Angel or Demon: Investigating the Plasticity Interventions' Impact on Backdoor Threats in Deep Reinforcement Learning
by: Ma, Oubo, et al.
Published: (2026)
by: Ma, Oubo, et al.
Published: (2026)
Multi-Agent Reinforcement Learning for Assessing False-Data Injection Attacks on Transportation Networks
by: Eghtesad, Taha, et al.
Published: (2023)
by: Eghtesad, Taha, et al.
Published: (2023)
Similar Items
-
Evading Deep Learning-Based Malware Detectors via Obfuscation: A Deep Reinforcement Learning Approach
by: Etter, Brian, et al.
Published: (2024) -
Multi-Agent Penetration Testing AI for the Web
by: David, Isaac, et al.
Published: (2025) -
StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors
by: Ranganath, Suraj, et al.
Published: (2026) -
Patch2Vuln: Agentic Reconstruction of Vulnerabilities from Linux Distribution Binary Patches
by: David, Isaac, et al.
Published: (2026) -
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications
by: David, Isaac, et al.
Published: (2026)