DiffuseDef: Improved Robustness to Adversarial Attacks via Iterative Denoising
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Zhenhao, Zhou, Huichi, Rei, Marek, Specia, Lucia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MPAT: Building Robust Deep Neural Networks against Textual Adversarial Attacks
di: Zhang, Fangyuan, et al.
Pubblicazione: (2024)
di: Zhang, Fangyuan, et al.
Pubblicazione: (2024)
MoESD: Mixture of Experts Stable Diffusion to Mitigate Gender Bias
di: Wang, Guorun, et al.
Pubblicazione: (2024)
di: Wang, Guorun, et al.
Pubblicazione: (2024)
From Understanding to Utilization: A Survey on Explainability for Large Language Models
di: Luo, Haoyan, et al.
Pubblicazione: (2024)
di: Luo, Haoyan, et al.
Pubblicazione: (2024)
Tuning Language Models by Mixture-of-Depths Ensemble
di: Luo, Haoyan, et al.
Pubblicazione: (2024)
di: Luo, Haoyan, et al.
Pubblicazione: (2024)
Evaluate-and-Purify: Fortifying Code Language Models Against Adversarial Attacks Using LLM-as-a-Judge
di: Mu, Wenhan, et al.
Pubblicazione: (2025)
di: Mu, Wenhan, et al.
Pubblicazione: (2025)
Meta-Reasoning Improves Tool Use in Large Language Models
di: Alazraki, Lisa, et al.
Pubblicazione: (2024)
di: Alazraki, Lisa, et al.
Pubblicazione: (2024)
SemanticShield: LLM-Powered Audits Expose Shilling Attacks in Recommender Systems
di: Li, Kaihong, et al.
Pubblicazione: (2025)
di: Li, Kaihong, et al.
Pubblicazione: (2025)
TrustRAG: Enhancing Robustness and Trustworthiness in Retrieval-Augmented Generation
di: Zhou, Huichi, et al.
Pubblicazione: (2025)
di: Zhou, Huichi, et al.
Pubblicazione: (2025)
Distilling Robustness into Natural Language Inference Models with Domain-Targeted Augmentation
di: Stacey, Joe, et al.
Pubblicazione: (2023)
di: Stacey, Joe, et al.
Pubblicazione: (2023)
Discourse Features Enhance Detection of Document-Level Machine-Generated Content
di: Li, Yupei, et al.
Pubblicazione: (2024)
di: Li, Yupei, et al.
Pubblicazione: (2024)
Fine-tuning with RAG for Improving LLM Learning of New Skills
di: Ibrahim, Humaid, et al.
Pubblicazione: (2025)
di: Ibrahim, Humaid, et al.
Pubblicazione: (2025)
Enhancing LLM Robustness to Perturbed Instructions: An Empirical Study
di: Agrawal, Aryan, et al.
Pubblicazione: (2025)
di: Agrawal, Aryan, et al.
Pubblicazione: (2025)
StateAct: Enhancing LLM Base Agents via Self-prompting and State-tracking
di: Rozanov, Nikolai, et al.
Pubblicazione: (2024)
di: Rozanov, Nikolai, et al.
Pubblicazione: (2024)
Adversarial Attack for Explanation Robustness of Rationalization Models
di: Zhang, Yuankai, et al.
Pubblicazione: (2024)
di: Zhang, Yuankai, et al.
Pubblicazione: (2024)
On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks
di: Obadinma, Stephen, et al.
Pubblicazione: (2025)
di: Obadinma, Stephen, et al.
Pubblicazione: (2025)
TabRAG: Improving Tabular Document Question Answering for Retrieval Augmented Generation via Structured Representations
di: Si, Jacob, et al.
Pubblicazione: (2025)
di: Si, Jacob, et al.
Pubblicazione: (2025)
Robustness of Large Language Models Against Adversarial Attacks
di: Tao, Yiyi, et al.
Pubblicazione: (2024)
di: Tao, Yiyi, et al.
Pubblicazione: (2024)
DefVerify: Do Hate Speech Models Reflect Their Dataset's Definition?
di: Khurana, Urja, et al.
Pubblicazione: (2024)
di: Khurana, Urja, et al.
Pubblicazione: (2024)
Def-DTS: Deductive Reasoning for Open-domain Dialogue Topic Segmentation
di: Lee, Seungmin, et al.
Pubblicazione: (2025)
di: Lee, Seungmin, et al.
Pubblicazione: (2025)
LinguIUTics at PsyDefDetect: Iterative Imbalance-Aware Fine-tuning of Qwen3-8B for Psychological Defense Mechanism Classification
di: Adib, Shefayat E Shams, et al.
Pubblicazione: (2026)
di: Adib, Shefayat E Shams, et al.
Pubblicazione: (2026)
Unpacking Robustness in Inflectional Languages: Adversarial Evaluation and Mechanistic Insights
di: Walkowiak, Paweł, et al.
Pubblicazione: (2025)
di: Walkowiak, Paweł, et al.
Pubblicazione: (2025)
Robustness of Misinformation Classification Systems to Adversarial Examples Through BeamAttack
di: Fazla, Arnisa, et al.
Pubblicazione: (2025)
di: Fazla, Arnisa, et al.
Pubblicazione: (2025)
DefAn: Definitive Answer Dataset for LLMs Hallucination Evaluation
di: Rahman, A B M Ashikur, et al.
Pubblicazione: (2024)
di: Rahman, A B M Ashikur, et al.
Pubblicazione: (2024)
Moral Reasoning Across Languages: The Critical Role of Low-Resource Languages in LLMs
di: Zhou, Huichi, et al.
Pubblicazione: (2025)
di: Zhou, Huichi, et al.
Pubblicazione: (2025)
Improving the OOD Performance of Closed-Source LLMs on NLI Through Strategic Data Selection
di: Stacey, Joe, et al.
Pubblicazione: (2025)
di: Stacey, Joe, et al.
Pubblicazione: (2025)
SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It)
di: Meeus, Matthieu, et al.
Pubblicazione: (2024)
di: Meeus, Matthieu, et al.
Pubblicazione: (2024)
Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
TF-Attack: Transferable and Fast Adversarial Attacks on Large Language Models
di: Li, Zelin, et al.
Pubblicazione: (2024)
di: Li, Zelin, et al.
Pubblicazione: (2024)
NLPerturbator: Studying the Robustness of Code LLMs to Natural Language Variations
di: Chen, Junkai, et al.
Pubblicazione: (2024)
di: Chen, Junkai, et al.
Pubblicazione: (2024)
RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation
di: Li, Tianjiao, et al.
Pubblicazione: (2025)
di: Li, Tianjiao, et al.
Pubblicazione: (2025)
SciDef: Automating Definition Extraction from Academic Literature with Large Language Models
di: Kučera, Filip, et al.
Pubblicazione: (2026)
di: Kučera, Filip, et al.
Pubblicazione: (2026)
Self Iterative Label Refinement via Robust Unlabeled Learning
di: Asano, Hikaru, et al.
Pubblicazione: (2025)
di: Asano, Hikaru, et al.
Pubblicazione: (2025)
Robust Vision-Language Models via Tensor Decomposition: A Defense Against Adversarial Attacks
di: Patel, Het, et al.
Pubblicazione: (2025)
di: Patel, Het, et al.
Pubblicazione: (2025)
Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
di: Raina, Vyas, et al.
Pubblicazione: (2024)
di: Raina, Vyas, et al.
Pubblicazione: (2024)
Robust Fake News Detection using Large Language Models under Adversarial Sentiment Attacks
di: Tahmasebi, Sahar, et al.
Pubblicazione: (2026)
di: Tahmasebi, Sahar, et al.
Pubblicazione: (2026)
REAL-IoT: Characterizing GNN Intrusion Detection Robustness under Practical Adversarial Attack
di: Zhan, Zhonghao, et al.
Pubblicazione: (2025)
di: Zhan, Zhonghao, et al.
Pubblicazione: (2025)
The Best Defense is Attack: Repairing Semantics in Textual Adversarial Examples
di: Yang, Heng, et al.
Pubblicazione: (2023)
di: Yang, Heng, et al.
Pubblicazione: (2023)
AIPO: Improving Training Objective for Iterative Preference Optimization
di: Shen, Yaojie, et al.
Pubblicazione: (2024)
di: Shen, Yaojie, et al.
Pubblicazione: (2024)
Camouflage is all you need: Evaluating and Enhancing Language Model Robustness Against Camouflage Adversarial Attacks
di: Huertas-García, Álvaro, et al.
Pubblicazione: (2024)
di: Huertas-García, Álvaro, et al.
Pubblicazione: (2024)
Time-To-Inconsistency: A Survival Analysis of Large Language Model Robustness to Adversarial Attacks
di: Li, Yubo, et al.
Pubblicazione: (2025)
di: Li, Yubo, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MPAT: Building Robust Deep Neural Networks against Textual Adversarial Attacks
di: Zhang, Fangyuan, et al.
Pubblicazione: (2024) -
MoESD: Mixture of Experts Stable Diffusion to Mitigate Gender Bias
di: Wang, Guorun, et al.
Pubblicazione: (2024) -
From Understanding to Utilization: A Survey on Explainability for Large Language Models
di: Luo, Haoyan, et al.
Pubblicazione: (2024) -
Tuning Language Models by Mixture-of-Depths Ensemble
di: Luo, Haoyan, et al.
Pubblicazione: (2024) -
Evaluate-and-Purify: Fortifying Code Language Models Against Adversarial Attacks Using LLM-as-a-Judge
di: Mu, Wenhan, et al.
Pubblicazione: (2025)