All You Need is "Leet": Evading Hate-speech Detection AI
Fuente:
arXiv
Guardado en:
| Autores principales: | Kahu, Sampanna Yashwant, Ahuja, Naman |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Fool Me If You Can: On the Robustness of Binary Code Similarity Detection Models against Semantics-preserving Transformations
por: Uhm, Jiyong, et al.
Publicado: (2026)
por: Uhm, Jiyong, et al.
Publicado: (2026)
Non-Linear Trajectory Modeling for Multi-Step Gradient Inversion Attacks in Federated Learning
por: Xia, Li, et al.
Publicado: (2025)
por: Xia, Li, et al.
Publicado: (2025)
MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization
por: Gu, Yongtong, et al.
Publicado: (2026)
por: Gu, Yongtong, et al.
Publicado: (2026)
Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models
por: Syed, Mohammed Sameer, et al.
Publicado: (2026)
por: Syed, Mohammed Sameer, et al.
Publicado: (2026)
Amplifying Training Data Exposure through Fine-Tuning with Pseudo-Labeled Memberships
por: Oh, Myung Gyo, et al.
Publicado: (2024)
por: Oh, Myung Gyo, et al.
Publicado: (2024)
I Know What You Said: Unveiling Hardware Cache Side-Channels in Local Large Language Model Inference
por: Gao, Zibo, et al.
Publicado: (2025)
por: Gao, Zibo, et al.
Publicado: (2025)
Activation Differences Reveal Backdoors: A Comparison of SAE Architectures
por: Kumar, Sachin
Publicado: (2026)
por: Kumar, Sachin
Publicado: (2026)
Dr. Jekyll and Mr. Hyde: Two Faces of LLMs
por: Collu, Matteo Gioele, et al.
Publicado: (2023)
por: Collu, Matteo Gioele, et al.
Publicado: (2023)
Cross-LLM Generalization of Behavioral Backdoor Detection in AI Agent Supply Chains
por: Sanna, Arun Chowdary
Publicado: (2025)
por: Sanna, Arun Chowdary
Publicado: (2025)
ADMIn: Attacks on Dataset, Model and Input. A Threat Model for AI Based Software
por: Kumar, Vimal, et al.
Publicado: (2024)
por: Kumar, Vimal, et al.
Publicado: (2024)
Fine-tuning RoBERTa for CVE-to-CWE Classification: A 125M Parameter Model Competitive with LLMs
por: Mosievskiy, Nikita
Publicado: (2026)
por: Mosievskiy, Nikita
Publicado: (2026)
Density-aware Sample-specific Attack
por: Wang, Qiyuan, et al.
Publicado: (2026)
por: Wang, Qiyuan, et al.
Publicado: (2026)
UnPII: Unlearning Personally Identifiable Information with Quantifiable Exposure Risk
por: Jeon, Intae, et al.
Publicado: (2026)
por: Jeon, Intae, et al.
Publicado: (2026)
SBASH: a Framework for Designing and Evaluating RAG vs. Prompt-Tuned LLM Honeypots
por: Adebimpe, Adetayo, et al.
Publicado: (2025)
por: Adebimpe, Adetayo, et al.
Publicado: (2025)
Code as a Weapon: A Consensus-Labeled Prompt Bank for Measuring Coding-Model Compliance with Malicious-Code Requests
por: Young, Richard J., et al.
Publicado: (2026)
por: Young, Richard J., et al.
Publicado: (2026)
Collaborative Zone-Adaptive Zero-Day Intrusion Detection for IoBT
por: Pasdar, Amirmohammad, et al.
Publicado: (2026)
por: Pasdar, Amirmohammad, et al.
Publicado: (2026)
A High-Recall Cost-Sensitive Machine Learning Framework for Real-Time Online Banking Transaction Fraud Detection
por: R., Karthikeyan V., et al.
Publicado: (2026)
por: R., Karthikeyan V., et al.
Publicado: (2026)
Context-Aware Web Attack Detection in Open-Source SIEM Systems via MITRE ATT&CK-Enriched Behavioral Profiling
por: Alboushy, Badr, et al.
Publicado: (2026)
por: Alboushy, Badr, et al.
Publicado: (2026)
RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry
por: Lv, Bo, et al.
Publicado: (2026)
por: Lv, Bo, et al.
Publicado: (2026)
LLM Detectors Still Fall Short of Real World: Case of LLM-Generated Short News-Like Posts
por: Gameiro, Henrique Da Silva, et al.
Publicado: (2024)
por: Gameiro, Henrique Da Silva, et al.
Publicado: (2024)
Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection
por: Lelle, Travis
Publicado: (2026)
por: Lelle, Travis
Publicado: (2026)
Optimisation of cyber insurance coverage with selection of cost effective security controls
por: Uuganbayar, Ganbayar, et al.
Publicado: (2025)
por: Uuganbayar, Ganbayar, et al.
Publicado: (2025)
Rubber Mallet: A Study of High Frequency Localized Bit Flips and Their Impact on Security
por: Adiletta, Andrew, et al.
Publicado: (2025)
por: Adiletta, Andrew, et al.
Publicado: (2025)
Active Authentication via Korean Keystrokes Under Varying LLM Assistance and Cognitive Contexts
por: Roh, Dong Hyun, et al.
Publicado: (2025)
por: Roh, Dong Hyun, et al.
Publicado: (2025)
Predicting known Vulnerabilities from Attack News: A Transformer-Based Approach
por: Othman, Refat, et al.
Publicado: (2026)
por: Othman, Refat, et al.
Publicado: (2026)
Enhancing Energy Sector Resilience: Integrating Security by Design Principles
por: Shirtz, Dov, et al.
Publicado: (2024)
por: Shirtz, Dov, et al.
Publicado: (2024)
Unveiling the Vulnerability of Private Fine-Tuning in Split-Based Frameworks for Large Language Models: A Bidirectionally Enhanced Attack
por: Chen, Guanzhong, et al.
Publicado: (2024)
por: Chen, Guanzhong, et al.
Publicado: (2024)
The Vehicle May Be Sick: Denial of Diagnostic Services by Exploiting the CAN Transport Protocol
por: Baek, Seungjin, et al.
Publicado: (2026)
por: Baek, Seungjin, et al.
Publicado: (2026)
A Relevance Model for Threat-Centric Ranking of Cybersecurity Vulnerabilities
por: McCoy, Corren, et al.
Publicado: (2024)
por: McCoy, Corren, et al.
Publicado: (2024)
Correlated-Sequence Differential Privacy
por: Luo, Yifan, et al.
Publicado: (2025)
por: Luo, Yifan, et al.
Publicado: (2025)
Memory Forensics Techniques for Automated Detection and Analysis of Go Malware
por: Ali, Hala, et al.
Publicado: (2026)
por: Ali, Hala, et al.
Publicado: (2026)
Phishing Detection System: An Ensemble Approach Using Character-Level CNN and Feature Engineering
por: Dubey, Rudra, et al.
Publicado: (2025)
por: Dubey, Rudra, et al.
Publicado: (2025)
Illuminating the Black Box: Real-Time Monitoring of Backdoor Unlearning in CNNs via Explainable AI
por: Hoang, Tien Dat
Publicado: (2025)
por: Hoang, Tien Dat
Publicado: (2025)
The Hiremath Early Detection (HED) Score: A Measure-Theoretic Evaluation Standard for Temporal Intelligence
por: Hiremath, Prakul Sunil
Publicado: (2026)
por: Hiremath, Prakul Sunil
Publicado: (2026)
EXHIB: A Benchmark for Realistic and Diverse Evaluation of Function Similarity in the Wild
por: Fan, Yiming, et al.
Publicado: (2026)
por: Fan, Yiming, et al.
Publicado: (2026)
Closing the Distribution Gap in Adversarial Training for LLMs
por: Hu, Chengzhi, et al.
Publicado: (2026)
por: Hu, Chengzhi, et al.
Publicado: (2026)
Multi-Agent Honeypot-Based Request-Response Context Dataset for Improved SQL Injection Detection Performance
por: Yu, Hao, et al.
Publicado: (2026)
por: Yu, Hao, et al.
Publicado: (2026)
Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security
por: Qi, Jinhu, et al.
Publicado: (2026)
por: Qi, Jinhu, et al.
Publicado: (2026)
Semantic Superiority vs. Forensic Efficiency: A Comparative Analysis of Deep Learning and Psycholinguistics for Business Email Compromise Detection
por: Adjei, Yaw Osei, et al.
Publicado: (2025)
por: Adjei, Yaw Osei, et al.
Publicado: (2025)
Reference-Free Spectral Analysis of EM Side-Channels for Always-on Hardware Trojan Detection
por: Tahghigh, Mahsa, et al.
Publicado: (2026)
por: Tahghigh, Mahsa, et al.
Publicado: (2026)
Ejemplares similares
-
Fool Me If You Can: On the Robustness of Binary Code Similarity Detection Models against Semantics-preserving Transformations
por: Uhm, Jiyong, et al.
Publicado: (2026) -
Non-Linear Trajectory Modeling for Multi-Step Gradient Inversion Attacks in Federated Learning
por: Xia, Li, et al.
Publicado: (2025) -
MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization
por: Gu, Yongtong, et al.
Publicado: (2026) -
Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models
por: Syed, Mohammed Sameer, et al.
Publicado: (2026) -
Amplifying Training Data Exposure through Fine-Tuning with Pseudo-Labeled Memberships
por: Oh, Myung Gyo, et al.
Publicado: (2024)