CERT-ED: Certifiably Robust Text Classification for Edit Distance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Zhuoqun, Marchant, Neil G, Ohrimenko, Olga, Rubinstein, Benjamin I. P. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AdaptDel: Adaptable Deletion Rate Randomized Smoothing for Certified Robustness
von: Huang, Zhuoqun, et al.
Veröffentlicht: (2025)
von: Huang, Zhuoqun, et al.
Veröffentlicht: (2025)
RS-Del: Edit Distance Robustness Certificates for Sequence Classifiers via Randomized Deletion
von: Huang, Zhuoqun, et al.
Veröffentlicht: (2023)
von: Huang, Zhuoqun, et al.
Veröffentlicht: (2023)
Getting a-Round Guarantees: Floating-Point Attacks on Certified Robustness
von: Jin, Jiankai, et al.
Veröffentlicht: (2022)
von: Jin, Jiankai, et al.
Veröffentlicht: (2022)
Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
von: Zhang, Xinyu, et al.
Veröffentlicht: (2023)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2023)
CR-UTP: Certified Robustness against Universal Text Perturbations on Large Language Models
von: Lou, Qian, et al.
Veröffentlicht: (2024)
von: Lou, Qian, et al.
Veröffentlicht: (2024)
Position: Certified Robustness Does Not (Yet) Imply Model Security
von: Cullen, Andrew C., et al.
Veröffentlicht: (2025)
von: Cullen, Andrew C., et al.
Veröffentlicht: (2025)
Certifiably Robust RAG against Retrieval Corruption
von: Xiang, Chong, et al.
Veröffentlicht: (2024)
von: Xiang, Chong, et al.
Veröffentlicht: (2024)
Information Leakage from Data Updates in Machine Learning Models
von: Hui, Tian, et al.
Veröffentlicht: (2023)
von: Hui, Tian, et al.
Veröffentlicht: (2023)
A Modified Word Saliency-Based Adversarial Attack on Text Classification Models
von: Waghela, Hetvi, et al.
Veröffentlicht: (2024)
von: Waghela, Hetvi, et al.
Veröffentlicht: (2024)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
How to Enhance Downstream Adversarial Robustness (almost) without Touching the Pre-Trained Foundation Model?
von: Liu, Meiqi, et al.
Veröffentlicht: (2025)
von: Liu, Meiqi, et al.
Veröffentlicht: (2025)
Are We There Yet? Timing and Floating-Point Attacks on Differential Privacy Systems
von: Jin, Jiankai, et al.
Veröffentlicht: (2021)
von: Jin, Jiankai, et al.
Veröffentlicht: (2021)
Watermarks Attack Watermarks: Re-Watermarking as a Generic Removal Strategy
von: Bulychev, Maria, et al.
Veröffentlicht: (2026)
von: Bulychev, Maria, et al.
Veröffentlicht: (2026)
Certifiably Robust Image Watermark
von: Jiang, Zhengyuan, et al.
Veröffentlicht: (2024)
von: Jiang, Zhengyuan, et al.
Veröffentlicht: (2024)
Network Traffic Classification Using Machine Learning, Transformer, and Large Language Models
von: Antari, Ahmad, et al.
Veröffentlicht: (2025)
von: Antari, Ahmad, et al.
Veröffentlicht: (2025)
Malware Classification from Memory Dumps Using Machine Learning, Transformers, and Large Language Models
von: Dweib, Areej, et al.
Veröffentlicht: (2025)
von: Dweib, Areej, et al.
Veröffentlicht: (2025)
Certifying LLM Safety against Adversarial Prompting
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
Transferable Embedding Inversion Attack: Uncovering Privacy Risks in Text Embeddings without Model Queries
von: Huang, Yu-Hsiang, et al.
Veröffentlicht: (2024)
von: Huang, Yu-Hsiang, et al.
Veröffentlicht: (2024)
A Robust Cybersecurity Topic Classification Tool
von: Pelofske, Elijah, et al.
Veröffentlicht: (2021)
von: Pelofske, Elijah, et al.
Veröffentlicht: (2021)
Graded Suspiciousness of Adversarial Texts to Human
von: Tonni, Shakila Mahjabin, et al.
Veröffentlicht: (2024)
von: Tonni, Shakila Mahjabin, et al.
Veröffentlicht: (2024)
A Transfer Attack to Image Watermarks
von: Hu, Yuepeng, et al.
Veröffentlicht: (2024)
von: Hu, Yuepeng, et al.
Veröffentlicht: (2024)
Revisiting the Robustness of Watermarking to Paraphrasing Attacks
von: Rastogi, Saksham, et al.
Veröffentlicht: (2024)
von: Rastogi, Saksham, et al.
Veröffentlicht: (2024)
Differentially Private Knowledge Distillation via Synthetic Text Generation
von: Flemings, James, et al.
Veröffentlicht: (2024)
von: Flemings, James, et al.
Veröffentlicht: (2024)
Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs
von: Mathew, Yohan, et al.
Veröffentlicht: (2024)
von: Mathew, Yohan, et al.
Veröffentlicht: (2024)
Deep Active Learning with Crowdsourcing Data for Privacy Policy Classification
von: Qiu, Wenjun, et al.
Veröffentlicht: (2020)
von: Qiu, Wenjun, et al.
Veröffentlicht: (2020)
Semantic Stealth: Adversarial Text Attacks on NLP Using Several Methods
von: Dey, Roopkatha, et al.
Veröffentlicht: (2024)
von: Dey, Roopkatha, et al.
Veröffentlicht: (2024)
Differentially Private Synthetic Text Generation for Retrieval-Augmented Generation (RAG)
von: Mori, Junki, et al.
Veröffentlicht: (2025)
von: Mori, Junki, et al.
Veröffentlicht: (2025)
ContinuousBench: Can Differentially Private Synthetic Text Improve Capabilities?
von: Liu, Peihan, et al.
Veröffentlicht: (2026)
von: Liu, Peihan, et al.
Veröffentlicht: (2026)
Enhancing Adversarial Text Attacks on BERT Models with Projected Gradient Descent
von: Waghela, Hetvi, et al.
Veröffentlicht: (2024)
von: Waghela, Hetvi, et al.
Veröffentlicht: (2024)
Reformulation is All You Need: Addressing Malicious Text Features in DNNs
von: Jiang, Yi, et al.
Veröffentlicht: (2025)
von: Jiang, Yi, et al.
Veröffentlicht: (2025)
The Canary's Echo: Auditing Privacy Risks of LLM-Generated Synthetic Text
von: Meeus, Matthieu, et al.
Veröffentlicht: (2025)
von: Meeus, Matthieu, et al.
Veröffentlicht: (2025)
TextSeal: A Localized LLM Watermark for Provenance & Distillation Protection
von: Sander, Tom, et al.
Veröffentlicht: (2026)
von: Sander, Tom, et al.
Veröffentlicht: (2026)
Robust Distortion-free Watermarks for Language Models
von: Kuditipudi, Rohith, et al.
Veröffentlicht: (2023)
von: Kuditipudi, Rohith, et al.
Veröffentlicht: (2023)
MIA-Tuner: Adapting Large Language Models as Pre-training Text Detector
von: Fu, Wenjie, et al.
Veröffentlicht: (2024)
von: Fu, Wenjie, et al.
Veröffentlicht: (2024)
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
von: Naseh, Ali, et al.
Veröffentlicht: (2025)
von: Naseh, Ali, et al.
Veröffentlicht: (2025)
InvisibleInk: High-Utility and Low-Cost Text Generation with Differential Privacy
von: Vinod, Vishnu, et al.
Veröffentlicht: (2025)
von: Vinod, Vishnu, et al.
Veröffentlicht: (2025)
Edit Distance Robust Watermarks via Indexing Pseudorandom Codes
von: Golowich, Noah, et al.
Veröffentlicht: (2024)
von: Golowich, Noah, et al.
Veröffentlicht: (2024)
Two Birds with One Stone: Multi-Task Detection and Attribution of LLM-Generated Text
von: Rao, Zixin, et al.
Veröffentlicht: (2025)
von: Rao, Zixin, et al.
Veröffentlicht: (2025)
Humanizing Machine-Generated Content: Evading AI-Text Detection through Adversarial Attack
von: Zhou, Ying, et al.
Veröffentlicht: (2024)
von: Zhou, Ying, et al.
Veröffentlicht: (2024)
FreqMark: Frequency-Based Watermark for Sentence-Level Detection of LLM-Generated Text
von: Xu, Zhenyu, et al.
Veröffentlicht: (2024)
von: Xu, Zhenyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AdaptDel: Adaptable Deletion Rate Randomized Smoothing for Certified Robustness
von: Huang, Zhuoqun, et al.
Veröffentlicht: (2025) -
RS-Del: Edit Distance Robustness Certificates for Sequence Classifiers via Randomized Deletion
von: Huang, Zhuoqun, et al.
Veröffentlicht: (2023) -
Getting a-Round Guarantees: Floating-Point Attacks on Certified Robustness
von: Jin, Jiankai, et al.
Veröffentlicht: (2022) -
Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
von: Zhang, Xinyu, et al.
Veröffentlicht: (2023) -
CR-UTP: Certified Robustness against Universal Text Perturbations on Large Language Models
von: Lou, Qian, et al.
Veröffentlicht: (2024)