DINA: A Dual Defense Framework Against Internal Noise and External Attacks in Natural Language Processing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chuang, Ko-Wei, Huang, Hen-Hsen, Li, Tsai-Yen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915434279731200
author Chuang, Ko-Wei
Huang, Hen-Hsen
Li, Tsai-Yen
author_facet Chuang, Ko-Wei
Huang, Hen-Hsen
Li, Tsai-Yen
contents As large language models (LLMs) and generative AI become increasingly integrated into customer service and moderation applications, adversarial threats emerge from both external manipulations and internal label corruption. In this work, we identify and systematically address these dual adversarial threats by introducing DINA (Dual Defense Against Internal Noise and Adversarial Attacks), a novel unified framework tailored specifically for NLP. Our approach adapts advanced noisy-label learning methods from computer vision and integrates them with adversarial training to simultaneously mitigate internal label sabotage and external adversarial perturbations. Extensive experiments conducted on a real-world dataset from an online gaming service demonstrate that DINA significantly improves model robustness and accuracy compared to baseline models. Our findings not only highlight the critical necessity of dual-threat defenses but also offer practical strategies for safeguarding NLP systems in realistic adversarial scenarios, underscoring broader implications for fair and responsible AI deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05671
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DINA: A Dual Defense Framework Against Internal Noise and External Attacks in Natural Language Processing
Chuang, Ko-Wei
Huang, Hen-Hsen
Li, Tsai-Yen
Cryptography and Security
Computation and Language
As large language models (LLMs) and generative AI become increasingly integrated into customer service and moderation applications, adversarial threats emerge from both external manipulations and internal label corruption. In this work, we identify and systematically address these dual adversarial threats by introducing DINA (Dual Defense Against Internal Noise and Adversarial Attacks), a novel unified framework tailored specifically for NLP. Our approach adapts advanced noisy-label learning methods from computer vision and integrates them with adversarial training to simultaneously mitigate internal label sabotage and external adversarial perturbations. Extensive experiments conducted on a real-world dataset from an online gaming service demonstrate that DINA significantly improves model robustness and accuracy compared to baseline models. Our findings not only highlight the critical necessity of dual-threat defenses but also offer practical strategies for safeguarding NLP systems in realistic adversarial scenarios, underscoring broader implications for fair and responsible AI deployment.
title DINA: A Dual Defense Framework Against Internal Noise and External Attacks in Natural Language Processing
topic Cryptography and Security
Computation and Language
url https://arxiv.org/abs/2508.05671