Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework
Fuente:
arXiv
Guardado en:
| Autores principales: | Pisano, Matthew, Ly, Peter, Sanders, Abraham, Yao, Bingsheng, Wang, Dakuo, Strzalkowski, Tomek, Si, Mei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multi-Granularity Tibetan Textual Adversarial Attack Method Based on Masked Language Model
por: Cao, Xi, et al.
Publicado: (2024)
por: Cao, Xi, et al.
Publicado: (2024)
Autonomous Chain-of-Thought Distillation for Graph-Based Fraud Detection
por: Li, Yuan, et al.
Publicado: (2026)
por: Li, Yuan, et al.
Publicado: (2026)
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
por: Maloyan, Narek, et al.
Publicado: (2025)
por: Maloyan, Narek, et al.
Publicado: (2025)
Semantic-Preserving Adversarial Attacks on LLMs: An Adaptive Greedy Binary Search Approach
por: Zhang, Chong, et al.
Publicado: (2025)
por: Zhang, Chong, et al.
Publicado: (2025)
Adversarial Robustness through Dynamic Ensemble Learning
por: Waghela, Hetvi, et al.
Publicado: (2024)
por: Waghela, Hetvi, et al.
Publicado: (2024)
Humanizing Machine-Generated Content: Evading AI-Text Detection through Adversarial Attack
por: Zhou, Ying, et al.
Publicado: (2024)
por: Zhou, Ying, et al.
Publicado: (2024)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
por: Wang, Jiongxiao, et al.
Publicado: (2024)
por: Wang, Jiongxiao, et al.
Publicado: (2024)
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
por: Li, Xiang, et al.
Publicado: (2025)
por: Li, Xiang, et al.
Publicado: (2025)
Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
por: Zhang, Xinyu, et al.
Publicado: (2023)
por: Zhang, Xinyu, et al.
Publicado: (2023)
Finding a Wolf in Sheep's Clothing: Combating Adversarial Text-To-Image Prompts with Text Summarization
por: Cooper, Portia, et al.
Publicado: (2024)
por: Cooper, Portia, et al.
Publicado: (2024)
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance
por: Huang, Caishuang, et al.
Publicado: (2024)
por: Huang, Caishuang, et al.
Publicado: (2024)
Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial Training
por: Li, Yuanfan, et al.
Publicado: (2025)
por: Li, Yuanfan, et al.
Publicado: (2025)
A Modified Word Saliency-Based Adversarial Attack on Text Classification Models
por: Waghela, Hetvi, et al.
Publicado: (2024)
por: Waghela, Hetvi, et al.
Publicado: (2024)
Pay Attention to the Robustness of Chinese Minority Language Models! Syllable-level Textual Adversarial Attack on Tibetan Script
por: Cao, Xi, et al.
Publicado: (2024)
por: Cao, Xi, et al.
Publicado: (2024)
Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
por: Wei, Zhang, et al.
Publicado: (2025)
por: Wei, Zhang, et al.
Publicado: (2025)
Emoti-Attack: Zero-Perturbation Adversarial Attacks on NLP Systems via Emoji Sequences
por: Zhang, Yangshijie
Publicado: (2025)
por: Zhang, Yangshijie
Publicado: (2025)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
por: Liu, Fan, et al.
Publicado: (2024)
por: Liu, Fan, et al.
Publicado: (2024)
IDT: Dual-Task Adversarial Attacks for Privacy Protection
por: Faustini, Pedro, et al.
Publicado: (2024)
por: Faustini, Pedro, et al.
Publicado: (2024)
Towards More Realistic Extraction Attacks: An Adversarial Perspective
por: More, Yash, et al.
Publicado: (2024)
por: More, Yash, et al.
Publicado: (2024)
Black-Box Guardrail Reverse-engineering Attack
por: Yao, Hongwei, et al.
Publicado: (2025)
por: Yao, Hongwei, et al.
Publicado: (2025)
Enhance Robustness of Language Models Against Variation Attack through Graph Integration
por: Xiong, Zi, et al.
Publicado: (2024)
por: Xiong, Zi, et al.
Publicado: (2024)
Task-Agnostic Detector for Insertion-Based Backdoor Attacks
por: Lyu, Weimin, et al.
Publicado: (2024)
por: Lyu, Weimin, et al.
Publicado: (2024)
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
por: Brown, Hannah, et al.
Publicado: (2024)
por: Brown, Hannah, et al.
Publicado: (2024)
Fast Adversarial Attacks on Language Models In One GPU Minute
por: Sadasivan, Vinu Sankar, et al.
Publicado: (2024)
por: Sadasivan, Vinu Sankar, et al.
Publicado: (2024)
Adversarial Attacks Against Automated Fact-Checking: A Survey
por: Liu, Fanzhen, et al.
Publicado: (2025)
por: Liu, Fanzhen, et al.
Publicado: (2025)
Look Twice before You Leap: A Rational Framework for Localized Adversarial Anonymization
por: Duan, Donghang, et al.
Publicado: (2025)
por: Duan, Donghang, et al.
Publicado: (2025)
Attacks against Abstractive Text Summarization Models through Lead Bias and Influence Functions
por: Thota, Poojitha, et al.
Publicado: (2024)
por: Thota, Poojitha, et al.
Publicado: (2024)
Black-Box Adversarial Attacks on LLM-Based Code Completion
por: Jenko, Slobodan, et al.
Publicado: (2024)
por: Jenko, Slobodan, et al.
Publicado: (2024)
Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space
por: Huang, Yao, et al.
Publicado: (2025)
por: Huang, Yao, et al.
Publicado: (2025)
Chain-of-Lure: A Universal Jailbreak Attack Framework using Unconstrained Synthetic Narratives
por: Chang, Wenhan, et al.
Publicado: (2025)
por: Chang, Wenhan, et al.
Publicado: (2025)
Adversarial Attack on Large Language Models using Exponentiated Gradient Descent
por: Biswas, Sajib, et al.
Publicado: (2025)
por: Biswas, Sajib, et al.
Publicado: (2025)
Semantic Stealth: Adversarial Text Attacks on NLP Using Several Methods
por: Dey, Roopkatha, et al.
Publicado: (2024)
por: Dey, Roopkatha, et al.
Publicado: (2024)
Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation
por: Li, Qizhang, et al.
Publicado: (2024)
por: Li, Qizhang, et al.
Publicado: (2024)
Token-Modification Adversarial Attacks for Natural Language Processing: A Survey
por: Roth, Tom, et al.
Publicado: (2021)
por: Roth, Tom, et al.
Publicado: (2021)
Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations
por: Teja, Lekkala Sai, et al.
Publicado: (2025)
por: Teja, Lekkala Sai, et al.
Publicado: (2025)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
por: Mu, Junjie, et al.
Publicado: (2025)
por: Mu, Junjie, et al.
Publicado: (2025)
Enhancing Adversarial Text Attacks on BERT Models with Projected Gradient Descent
por: Waghela, Hetvi, et al.
Publicado: (2024)
por: Waghela, Hetvi, et al.
Publicado: (2024)
The Communication-Friendly Privacy-Preserving Machine Learning against Malicious Adversaries
por: Lu, Tianpei, et al.
Publicado: (2024)
por: Lu, Tianpei, et al.
Publicado: (2024)
RTD-Guard: A Black-Box Textual Adversarial Detection Framework via Replacement Token Detection
por: Zhu, He, et al.
Publicado: (2026)
por: Zhu, He, et al.
Publicado: (2026)
DINA: A Dual Defense Framework Against Internal Noise and External Attacks in Natural Language Processing
por: Chuang, Ko-Wei, et al.
Publicado: (2025)
por: Chuang, Ko-Wei, et al.
Publicado: (2025)
Ejemplares similares
-
Multi-Granularity Tibetan Textual Adversarial Attack Method Based on Masked Language Model
por: Cao, Xi, et al.
Publicado: (2024) -
Autonomous Chain-of-Thought Distillation for Graph-Based Fraud Detection
por: Li, Yuan, et al.
Publicado: (2026) -
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
por: Maloyan, Narek, et al.
Publicado: (2025) -
Semantic-Preserving Adversarial Attacks on LLMs: An Adaptive Greedy Binary Search Approach
por: Zhang, Chong, et al.
Publicado: (2025) -
Adversarial Robustness through Dynamic Ensemble Learning
por: Waghela, Hetvi, et al.
Publicado: (2024)