CR-UTP: Certified Robustness against Universal Text Perturbations on Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lou, Qian, Liang, Xin, Xue, Jiaqi, Zhang, Yancheng, Xie, Rui, Zheng, Mengxin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RobPI: Robust Private Inference against Malicious Client
von: Xue, Jiaqi, et al.
Veröffentlicht: (2026)
von: Xue, Jiaqi, et al.
Veröffentlicht: (2026)
BadFair: Backdoored Fairness Attacks with Group-conditioned Triggers
von: Xue, Jiaqi, et al.
Veröffentlicht: (2024)
von: Xue, Jiaqi, et al.
Veröffentlicht: (2024)
BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models
von: Xue, Jiaqi, et al.
Veröffentlicht: (2024)
von: Xue, Jiaqi, et al.
Veröffentlicht: (2024)
CipherPrune: Efficient and Scalable Private Transformer Inference
von: Zhang, Yancheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yancheng, et al.
Veröffentlicht: (2025)
Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Texts
von: Hastuti, Rochana Prih, et al.
Veröffentlicht: (2025)
von: Hastuti, Rochana Prih, et al.
Veröffentlicht: (2025)
TFHE-Coder: Evaluating LLM-agentic Fully Homomorphic Encryption Code Generation
von: Kumar, Mayank, et al.
Veröffentlicht: (2025)
von: Kumar, Mayank, et al.
Veröffentlicht: (2025)
PRO: Enabling Precise and Robust Text Watermark for Open-Source LLMs
von: Xue, Jiaqi, et al.
Veröffentlicht: (2025)
von: Xue, Jiaqi, et al.
Veröffentlicht: (2025)
Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
von: Zhang, Xinyu, et al.
Veröffentlicht: (2023)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2023)
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
von: Mangaokar, Neal, et al.
Veröffentlicht: (2024)
von: Mangaokar, Neal, et al.
Veröffentlicht: (2024)
Certifiably Robust RAG against Retrieval Corruption
von: Xiang, Chong, et al.
Veröffentlicht: (2024)
von: Xiang, Chong, et al.
Veröffentlicht: (2024)
Collective Certified Robustness against Graph Injection Attacks
von: Lai, Yuni, et al.
Veröffentlicht: (2024)
von: Lai, Yuni, et al.
Veröffentlicht: (2024)
Towards Safe AI Clinicians: A Comprehensive Study on Large Language Model Jailbreaking in Healthcare
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
Securing Transformer-based AI Execution via Unified TEEs and Crypto-protected Accelerators
von: Xue, Jiaqi, et al.
Veröffentlicht: (2025)
von: Xue, Jiaqi, et al.
Veröffentlicht: (2025)
CERT-ED: Certifiably Robust Text Classification for Edit Distance
von: Huang, Zhuoqun, et al.
Veröffentlicht: (2024)
von: Huang, Zhuoqun, et al.
Veröffentlicht: (2024)
BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
LLMAtKGE: Large Language Models as Explainable Attackers against Knowledge Graph Embeddings
von: Li, Ting, et al.
Veröffentlicht: (2025)
von: Li, Ting, et al.
Veröffentlicht: (2025)
RobustKV: Defending Large Language Models against Jailbreak Attacks via KV Eviction
von: Jiang, Tanqiu, et al.
Veröffentlicht: (2024)
von: Jiang, Tanqiu, et al.
Veröffentlicht: (2024)
SecureRouter: Encrypted Routing for Efficient Secure Inference
von: Zhang, Yukuan, et al.
Veröffentlicht: (2026)
von: Zhang, Yukuan, et al.
Veröffentlicht: (2026)
zkVC: Fast Zero-Knowledge Proof for Private and Verifiable Computing
von: Zhang, Yancheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yancheng, et al.
Veröffentlicht: (2025)
A Certified Robust Watermark For Large Language Models
von: Feng, Xianheng, et al.
Veröffentlicht: (2024)
von: Feng, Xianheng, et al.
Veröffentlicht: (2024)
Imperceptible Jailbreaking against Large Language Models
von: Gao, Kuofeng, et al.
Veröffentlicht: (2025)
von: Gao, Kuofeng, et al.
Veröffentlicht: (2025)
Denial-of-Service Poisoning Attacks against Large Language Models
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
DictPFL: Efficient and Private Federated Learning on Encrypted Gradients
von: Xue, Jiaqi, et al.
Veröffentlicht: (2025)
von: Xue, Jiaqi, et al.
Veröffentlicht: (2025)
AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image Models
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
Probing the Robustness of Large Language Models Safety to Latent Perturbations
von: Gu, Tianle, et al.
Veröffentlicht: (2025)
von: Gu, Tianle, et al.
Veröffentlicht: (2025)
StructuralSleight: Automated Jailbreak Attacks on Large Language Models Utilizing Uncommon Text-Organization Structures
von: Li, Bangxin, et al.
Veröffentlicht: (2024)
von: Li, Bangxin, et al.
Veröffentlicht: (2024)
CryptoTrain: Fast Secure Training on Encrypted Dataset
von: Xue, Jiaqi, et al.
Veröffentlicht: (2024)
von: Xue, Jiaqi, et al.
Veröffentlicht: (2024)
Adversarial Text Generation with Dynamic Contextual Perturbation
von: Waghela, Hetvi, et al.
Veröffentlicht: (2025)
von: Waghela, Hetvi, et al.
Veröffentlicht: (2025)
Doubly-Universal Adversarial Perturbations: Deceiving Vision-Language Models Across Both Images and Text with a Single Perturbation
von: Kim, Hee-Seon, et al.
Veröffentlicht: (2024)
von: Kim, Hee-Seon, et al.
Veröffentlicht: (2024)
Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
von: Wei, Zhang, et al.
Veröffentlicht: (2025)
von: Wei, Zhang, et al.
Veröffentlicht: (2025)
Towards Label-Only Membership Inference Attack against Pre-trained Large Language Models
von: He, Yu, et al.
Veröffentlicht: (2025)
von: He, Yu, et al.
Veröffentlicht: (2025)
Majority Bit-Aware Watermarking For Large Language Models
von: Xu, Jiahao, et al.
Veröffentlicht: (2025)
von: Xu, Jiahao, et al.
Veröffentlicht: (2025)
RPP: A Certified Poisoned-Sample Detection Framework for Backdoor Attacks under Dataset Imbalance
von: Lin, Miao, et al.
Veröffentlicht: (2026)
von: Lin, Miao, et al.
Veröffentlicht: (2026)
DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack
von: Li, Hao, et al.
Veröffentlicht: (2025)
von: Li, Hao, et al.
Veröffentlicht: (2025)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
SoK: Can Fully Homomorphic Encryption Support General AI Computation? A Functional and Cost Analysis
von: Xue, Jiaqi, et al.
Veröffentlicht: (2025)
von: Xue, Jiaqi, et al.
Veröffentlicht: (2025)
Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language Models
von: He, Jiaming, et al.
Veröffentlicht: (2024)
von: He, Jiaming, et al.
Veröffentlicht: (2024)
Cross-Input Certified Training for Universal Perturbations
von: Xu, Changming, et al.
Veröffentlicht: (2024)
von: Xu, Changming, et al.
Veröffentlicht: (2024)
Privacy-Preserving Instructions for Aligning Large Language Models
von: Yu, Da, et al.
Veröffentlicht: (2024)
von: Yu, Da, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
RobPI: Robust Private Inference against Malicious Client
von: Xue, Jiaqi, et al.
Veröffentlicht: (2026) -
BadFair: Backdoored Fairness Attacks with Group-conditioned Triggers
von: Xue, Jiaqi, et al.
Veröffentlicht: (2024) -
BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models
von: Xue, Jiaqi, et al.
Veröffentlicht: (2024) -
CipherPrune: Efficient and Scalable Private Transformer Inference
von: Zhang, Yancheng, et al.
Veröffentlicht: (2025) -
Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Texts
von: Hastuti, Rochana Prih, et al.
Veröffentlicht: (2025)