Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Kaneko, Masahiro, Baldwin, Timothy |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
por: Brown, Hannah, et al.
Publicado: (2024)
por: Brown, Hannah, et al.
Publicado: (2024)
MPAT: Building Robust Deep Neural Networks against Textual Adversarial Attacks
por: Zhang, Fangyuan, et al.
Publicado: (2024)
por: Zhang, Fangyuan, et al.
Publicado: (2024)
Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
por: Zhang, Xinyu, et al.
Publicado: (2023)
por: Zhang, Xinyu, et al.
Publicado: (2023)
TFL: Targeted Bit-Flip Attack on Large Language Model
por: Guo, Jingkai, et al.
Publicado: (2026)
por: Guo, Jingkai, et al.
Publicado: (2026)
SBFA: Single Sneaky Bit Flip Attack to Break Large Language Models
por: Guo, Jingkai, et al.
Publicado: (2025)
por: Guo, Jingkai, et al.
Publicado: (2025)
IDT: Dual-Task Adversarial Attacks for Privacy Protection
por: Faustini, Pedro, et al.
Publicado: (2024)
por: Faustini, Pedro, et al.
Publicado: (2024)
Towards More Realistic Extraction Attacks: An Adversarial Perspective
por: More, Yash, et al.
Publicado: (2024)
por: More, Yash, et al.
Publicado: (2024)
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
por: Poppi, Samuele, et al.
Publicado: (2024)
por: Poppi, Samuele, et al.
Publicado: (2024)
Adversarial Attack on Large Language Models using Exponentiated Gradient Descent
por: Biswas, Sajib, et al.
Publicado: (2025)
por: Biswas, Sajib, et al.
Publicado: (2025)
Semantic Stealth: Adversarial Text Attacks on NLP Using Several Methods
por: Dey, Roopkatha, et al.
Publicado: (2024)
por: Dey, Roopkatha, et al.
Publicado: (2024)
Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation
por: Li, Qizhang, et al.
Publicado: (2024)
por: Li, Qizhang, et al.
Publicado: (2024)
Token-Modification Adversarial Attacks for Natural Language Processing: A Survey
por: Roth, Tom, et al.
Publicado: (2021)
por: Roth, Tom, et al.
Publicado: (2021)
Enhancing Adversarial Text Attacks on BERT Models with Projected Gradient Descent
por: Waghela, Hetvi, et al.
Publicado: (2024)
por: Waghela, Hetvi, et al.
Publicado: (2024)
Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks
por: Hu, Hanjiang, et al.
Publicado: (2025)
por: Hu, Hanjiang, et al.
Publicado: (2025)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
por: Zeng, Yifan, et al.
Publicado: (2024)
por: Zeng, Yifan, et al.
Publicado: (2024)
Transferable Embedding Inversion Attack: Uncovering Privacy Risks in Text Embeddings without Model Queries
por: Huang, Yu-Hsiang, et al.
Publicado: (2024)
por: Huang, Yu-Hsiang, et al.
Publicado: (2024)
A Modified Word Saliency-Based Adversarial Attack on Text Classification Models
por: Waghela, Hetvi, et al.
Publicado: (2024)
por: Waghela, Hetvi, et al.
Publicado: (2024)
Query-Based Adversarial Prompt Generation
por: Hayase, Jonathan, et al.
Publicado: (2024)
por: Hayase, Jonathan, et al.
Publicado: (2024)
MARAGE: Transferable Multi-Model Adversarial Attack for Retrieval-Augmented Generation Data Extraction
por: Hu, Xiao, et al.
Publicado: (2025)
por: Hu, Xiao, et al.
Publicado: (2025)
Humanizing Machine-Generated Content: Evading AI-Text Detection through Adversarial Attack
por: Zhou, Ying, et al.
Publicado: (2024)
por: Zhou, Ying, et al.
Publicado: (2024)
LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
por: Li, Ran, et al.
Publicado: (2025)
por: Li, Ran, et al.
Publicado: (2025)
DocMIA: Document-Level Membership Inference Attacks against DocVQA Models
por: Nguyen, Khanh, et al.
Publicado: (2025)
por: Nguyen, Khanh, et al.
Publicado: (2025)
Securing Large Language Models (LLMs) from Prompt Injection Attacks
por: Suri, Omar Farooq Khan, et al.
Publicado: (2025)
por: Suri, Omar Farooq Khan, et al.
Publicado: (2025)
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
por: Aldahoul, Nouar, et al.
Publicado: (2025)
por: Aldahoul, Nouar, et al.
Publicado: (2025)
SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It)
por: Meeus, Matthieu, et al.
Publicado: (2024)
por: Meeus, Matthieu, et al.
Publicado: (2024)
Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language Models
por: He, Jiaming, et al.
Publicado: (2024)
por: He, Jiaming, et al.
Publicado: (2024)
Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration
por: Fu, Wenjie, et al.
Publicado: (2023)
por: Fu, Wenjie, et al.
Publicado: (2023)
Certifying LLM Safety against Adversarial Prompting
por: Kumar, Aounon, et al.
Publicado: (2023)
por: Kumar, Aounon, et al.
Publicado: (2023)
BrainLeaks: On the Privacy-Preserving Properties of Neuromorphic Architectures against Model Inversion Attacks
por: Poursiami, Hamed, et al.
Publicado: (2024)
por: Poursiami, Hamed, et al.
Publicado: (2024)
Adversarial Decoding: Generating Readable Documents for Adversarial Objectives
por: Zhang, Collin, et al.
Publicado: (2024)
por: Zhang, Collin, et al.
Publicado: (2024)
The Resurgence of GCG Adversarial Attacks on Large Language Models
por: Tan, Yuting, et al.
Publicado: (2025)
por: Tan, Yuting, et al.
Publicado: (2025)
PrisonBreak: Jailbreaking Large Language Models with at Most Twenty-Five Targeted Bit-flips
por: Coalson, Zachary, et al.
Publicado: (2024)
por: Coalson, Zachary, et al.
Publicado: (2024)
Logits of API-Protected LLMs Leak Proprietary Information
por: Finlayson, Matthew, et al.
Publicado: (2024)
por: Finlayson, Matthew, et al.
Publicado: (2024)
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
por: Liang, Buyun, et al.
Publicado: (2026)
por: Liang, Buyun, et al.
Publicado: (2026)
HSF: Defending against Jailbreak Attacks with Hidden State Filtering
por: Qian, Cheng, et al.
Publicado: (2024)
por: Qian, Cheng, et al.
Publicado: (2024)
Adversarial Attacks on Parts of Speech: An Empirical Study in Text-to-Image Generation
por: Shahariar, G M, et al.
Publicado: (2024)
por: Shahariar, G M, et al.
Publicado: (2024)
On Evaluating The Performance of Watermarked Machine-Generated Texts Under Adversarial Attacks
por: Liu, Zesen, et al.
Publicado: (2024)
por: Liu, Zesen, et al.
Publicado: (2024)
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
por: Paulus, Anselm, et al.
Publicado: (2024)
por: Paulus, Anselm, et al.
Publicado: (2024)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
por: Chen, Taiye, et al.
Publicado: (2025)
por: Chen, Taiye, et al.
Publicado: (2025)
Graded Suspiciousness of Adversarial Texts to Human
por: Tonni, Shakila Mahjabin, et al.
Publicado: (2024)
por: Tonni, Shakila Mahjabin, et al.
Publicado: (2024)
Ejemplares similares
-
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
por: Brown, Hannah, et al.
Publicado: (2024) -
MPAT: Building Robust Deep Neural Networks against Textual Adversarial Attacks
por: Zhang, Fangyuan, et al.
Publicado: (2024) -
Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
por: Zhang, Xinyu, et al.
Publicado: (2023) -
TFL: Targeted Bit-Flip Attack on Large Language Model
por: Guo, Jingkai, et al.
Publicado: (2026) -
SBFA: Single Sneaky Bit Flip Attack to Break Large Language Models
por: Guo, Jingkai, et al.
Publicado: (2025)