Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Xinyu, Hong, Hanbin, Hong, Yuan, Huang, Peng, Wang, Binghui, Ba, Zhongjie, Ren, Kui |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Certifiable Black-Box Attacks with Randomized Adversarial Examples: Breaking Defenses with Provable Confidence
di: Hong, Hanbin, et al.
Pubblicazione: (2023)
di: Hong, Hanbin, et al.
Pubblicazione: (2023)
"Training robust watermarking model may hurt authentication!'' Exploring and Mitigating the Identity Leakage in Robust Watermarking
di: Zhang, Xinyu, et al.
Pubblicazione: (2026)
di: Zhang, Xinyu, et al.
Pubblicazione: (2026)
Towards Label-Only Membership Inference Attack against Pre-trained Large Language Models
di: He, Yu, et al.
Pubblicazione: (2025)
di: He, Yu, et al.
Pubblicazione: (2025)
Towards Strong Certified Defense with Universal Asymmetric Randomization
di: Hong, Hanbin, et al.
Pubblicazione: (2025)
di: Hong, Hanbin, et al.
Pubblicazione: (2025)
FedGMark: Certifiably Robust Watermarking for Federated Graph Learning
di: Yang, Yuxin, et al.
Pubblicazione: (2024)
di: Yang, Yuxin, et al.
Pubblicazione: (2024)
MPAT: Building Robust Deep Neural Networks against Textual Adversarial Attacks
di: Zhang, Fangyuan, et al.
Pubblicazione: (2024)
di: Zhang, Fangyuan, et al.
Pubblicazione: (2024)
Distributed Backdoor Attacks on Federated Graph Learning and Certified Defenses
di: Yang, Yuxin, et al.
Pubblicazione: (2024)
di: Yang, Yuxin, et al.
Pubblicazione: (2024)
Inf2Guard: An Information-Theoretic Framework for Learning Privacy-Preserving Representations against Inference Attacks
di: Noorbakhsh, Sayedeh Leila, et al.
Pubblicazione: (2024)
di: Noorbakhsh, Sayedeh Leila, et al.
Pubblicazione: (2024)
Robust Watermarks Leak: Channel-Aware Feature Extraction Enables Adversarial Watermark Manipulation
di: Ba, Zhongjie, et al.
Pubblicazione: (2025)
di: Ba, Zhongjie, et al.
Pubblicazione: (2025)
Practical, Generalizable and Robust Backdoor Attacks on Text-to-Image Diffusion Models
di: Dai, Haoran, et al.
Pubblicazione: (2025)
di: Dai, Haoran, et al.
Pubblicazione: (2025)
Pay Attention to the Robustness of Chinese Minority Language Models! Syllable-level Textual Adversarial Attack on Tibetan Script
di: Cao, Xi, et al.
Pubblicazione: (2024)
di: Cao, Xi, et al.
Pubblicazione: (2024)
Certifying Adapters: Enabling and Enhancing the Certification of Classifier Adversarial Robustness
di: Deng, Jieren, et al.
Pubblicazione: (2024)
di: Deng, Jieren, et al.
Pubblicazione: (2024)
ALIF: Low-Cost Adversarial Audio Attacks on Black-Box Speech Platforms using Linguistic Features
di: Cheng, Peng, et al.
Pubblicazione: (2024)
di: Cheng, Peng, et al.
Pubblicazione: (2024)
Multi-Granularity Tibetan Textual Adversarial Attack Method Based on Masked Language Model
di: Cao, Xi, et al.
Pubblicazione: (2024)
di: Cao, Xi, et al.
Pubblicazione: (2024)
CR-UTP: Certified Robustness against Universal Text Perturbations on Large Language Models
di: Lou, Qian, et al.
Pubblicazione: (2024)
di: Lou, Qian, et al.
Pubblicazione: (2024)
Certifiably Robust RAG against Retrieval Corruption
di: Xiang, Chong, et al.
Pubblicazione: (2024)
di: Xiang, Chong, et al.
Pubblicazione: (2024)
Collective Certified Robustness against Graph Injection Attacks
di: Lai, Yuni, et al.
Pubblicazione: (2024)
di: Lai, Yuni, et al.
Pubblicazione: (2024)
Data Extraction Attacks in Retrieval-Augmented Generation via Backdoors
di: Peng, Yuefeng, et al.
Pubblicazione: (2024)
di: Peng, Yuefeng, et al.
Pubblicazione: (2024)
RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
di: Xu, Huiyu, et al.
Pubblicazione: (2024)
di: Xu, Huiyu, et al.
Pubblicazione: (2024)
Certifying LLM Safety against Adversarial Prompting
di: Kumar, Aounon, et al.
Pubblicazione: (2023)
di: Kumar, Aounon, et al.
Pubblicazione: (2023)
CERT-ED: Certifiably Robust Text Classification for Edit Distance
di: Huang, Zhuoqun, et al.
Pubblicazione: (2024)
di: Huang, Zhuoqun, et al.
Pubblicazione: (2024)
RobustMask: Certified Robustness against Adversarial Neural Ranking Attack via Randomized Masking
di: Liu, Jiawei, et al.
Pubblicazione: (2025)
di: Liu, Jiawei, et al.
Pubblicazione: (2025)
A Learning-Based Attack Framework to Break SOTA Poisoning Defenses in Federated Learning
di: Yang, Yuxin, et al.
Pubblicazione: (2024)
di: Yang, Yuxin, et al.
Pubblicazione: (2024)
Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial Training
di: Li, Yuanfan, et al.
Pubblicazione: (2025)
di: Li, Yuanfan, et al.
Pubblicazione: (2025)
A Certified Robust Watermark For Large Language Models
di: Feng, Xianheng, et al.
Pubblicazione: (2024)
di: Feng, Xianheng, et al.
Pubblicazione: (2024)
Provably Robust Explainable Graph Neural Networks against Graph Perturbation Attacks
di: Li, Jiate, et al.
Pubblicazione: (2025)
di: Li, Jiate, et al.
Pubblicazione: (2025)
RTD-Guard: A Black-Box Textual Adversarial Detection Framework via Replacement Token Detection
di: Zhu, He, et al.
Pubblicazione: (2026)
di: Zhu, He, et al.
Pubblicazione: (2026)
Backdoor Attacks on Discrete Graph Diffusion Models
di: Wang, Jiawen, et al.
Pubblicazione: (2025)
di: Wang, Jiawen, et al.
Pubblicazione: (2025)
Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations
di: Teja, Lekkala Sai, et al.
Pubblicazione: (2025)
di: Teja, Lekkala Sai, et al.
Pubblicazione: (2025)
TSCheater: Generating High-Quality Tibetan Adversarial Texts via Visual Similarity
di: Cao, Xi, et al.
Pubblicazione: (2024)
di: Cao, Xi, et al.
Pubblicazione: (2024)
Learning Robust and Privacy-Preserving Representations via Information Theory
di: Zhang, Binghui, et al.
Pubblicazione: (2024)
di: Zhang, Binghui, et al.
Pubblicazione: (2024)
Fight Poison with Poison: Enhancing Robustness in Few-shot Machine-Generated Text Detection with Adversarial Training
di: Duan, Wenjing, et al.
Pubblicazione: (2026)
di: Duan, Wenjing, et al.
Pubblicazione: (2026)
Adversarial Text Generation with Dynamic Contextual Perturbation
di: Waghela, Hetvi, et al.
Pubblicazione: (2025)
di: Waghela, Hetvi, et al.
Pubblicazione: (2025)
Attacks against Abstractive Text Summarization Models through Lead Bias and Influence Functions
di: Thota, Poojitha, et al.
Pubblicazione: (2024)
di: Thota, Poojitha, et al.
Pubblicazione: (2024)
Universally Harmonizing Differential Privacy Mechanisms for Federated Learning: Boosting Accuracy and Convergence
di: Feng, Shuya, et al.
Pubblicazione: (2024)
di: Feng, Shuya, et al.
Pubblicazione: (2024)
Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information Optimization
di: Liu, Tiantian, et al.
Pubblicazione: (2024)
di: Liu, Tiantian, et al.
Pubblicazione: (2024)
Humanizing Machine-Generated Content: Evading AI-Text Detection through Adversarial Attack
di: Zhou, Ying, et al.
Pubblicazione: (2024)
di: Zhou, Ying, et al.
Pubblicazione: (2024)
Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks
di: Li, Ziqiang, et al.
Pubblicazione: (2024)
di: Li, Ziqiang, et al.
Pubblicazione: (2024)
The Communication-Friendly Privacy-Preserving Machine Learning against Malicious Adversaries
di: Lu, Tianpei, et al.
Pubblicazione: (2024)
di: Lu, Tianpei, et al.
Pubblicazione: (2024)
SurrogatePrompt: Bypassing the Safety Filter of Text-to-Image Models via Substitution
di: Ba, Zhongjie, et al.
Pubblicazione: (2023)
di: Ba, Zhongjie, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Certifiable Black-Box Attacks with Randomized Adversarial Examples: Breaking Defenses with Provable Confidence
di: Hong, Hanbin, et al.
Pubblicazione: (2023) -
"Training robust watermarking model may hurt authentication!'' Exploring and Mitigating the Identity Leakage in Robust Watermarking
di: Zhang, Xinyu, et al.
Pubblicazione: (2026) -
Towards Label-Only Membership Inference Attack against Pre-trained Large Language Models
di: He, Yu, et al.
Pubblicazione: (2025) -
Towards Strong Certified Defense with Universal Asymmetric Randomization
di: Hong, Hanbin, et al.
Pubblicazione: (2025) -
FedGMark: Certifiably Robust Watermarking for Federated Graph Learning
di: Yang, Yuxin, et al.
Pubblicazione: (2024)