Adversarial Defence without Adversarial Defence: Enhancing Language Model Robustness via Instance-level Principal Component Removal
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yang, Xiao, Chenghao, Li, Yizhi, Middleton, Stuart E., Moubayed, Noura Al, Lin, Chenghua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence Benchmarks
by: Wang, Yang, et al.
Published: (2025)
by: Wang, Yang, et al.
Published: (2025)
Audio Contrastive-based Fine-tuning: Decoupling Representation Learning and Classification
by: Wang, Yang, et al.
Published: (2023)
by: Wang, Yang, et al.
Published: (2023)
RAR-b: Reasoning as Retrieval Benchmark
by: Xiao, Chenghao, et al.
Published: (2024)
by: Xiao, Chenghao, et al.
Published: (2024)
X-ray Made Simple: Lay Radiology Report Generation and Robust Evaluation
by: Zhao, Kun, et al.
Published: (2024)
by: Zhao, Kun, et al.
Published: (2024)
Pixel Sentence Representation Learning
by: Xiao, Chenghao, et al.
Published: (2024)
by: Xiao, Chenghao, et al.
Published: (2024)
MuLD: The Multitask Long Document Benchmark
by: Hudson, G Thomas, et al.
Published: (2022)
by: Hudson, G Thomas, et al.
Published: (2022)
Early Detection and Reduction of Memorisation for Domain Adaptation and Instruction Tuning
by: Slack, Dean L., et al.
Published: (2025)
by: Slack, Dean L., et al.
Published: (2025)
Analyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal Representations
by: Xiao, Chenghao, et al.
Published: (2025)
by: Xiao, Chenghao, et al.
Published: (2025)
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
by: Wu, Siwei, et al.
Published: (2024)
by: Wu, Siwei, et al.
Published: (2024)
Interpreting Adversarial Attacks and Defences using Architectures with Enhanced Interpretability
by: Rao, Akshay G, et al.
Published: (2025)
by: Rao, Akshay G, et al.
Published: (2025)
Adaptive Randomized Smoothing: Certified Adversarial Robustness for Multi-Step Defences
by: Lyu, Saiyue, et al.
Published: (2024)
by: Lyu, Saiyue, et al.
Published: (2024)
Everything is a Video: Unifying Modalities through Next-Frame Prediction
by: Hudson, G. Thomas, et al.
Published: (2024)
by: Hudson, G. Thomas, et al.
Published: (2024)
Evaluating Defences against Unsafe Feedback in RLHF
by: Rosati, Domenic, et al.
Published: (2024)
by: Rosati, Domenic, et al.
Published: (2024)
On the Rigour of Scientific Writing: Criteria, Analysis, and Insights
by: James, Joseph, et al.
Published: (2024)
by: James, Joseph, et al.
Published: (2024)
Round Trip Translation Defence against Large Language Model Jailbreaking Attacks
by: Yung, Canaan, et al.
Published: (2024)
by: Yung, Canaan, et al.
Published: (2024)
LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents
by: Klem, Strahinja, et al.
Published: (2025)
by: Klem, Strahinja, et al.
Published: (2025)
Effective Distillation of Table-based Reasoning Ability from LLMs
by: Yang, Bohao, et al.
Published: (2023)
by: Yang, Bohao, et al.
Published: (2023)
Representation Noising: A Defence Mechanism Against Harmful Finetuning
by: Rosati, Domenic, et al.
Published: (2024)
by: Rosati, Domenic, et al.
Published: (2024)
Root Defence Strategies: Ensuring Safety of LLM at the Decoding Level
by: Zeng, Xinyi, et al.
Published: (2024)
by: Zeng, Xinyi, et al.
Published: (2024)
MIEB: Massive Image Embedding Benchmark
by: Xiao, Chenghao, et al.
Published: (2025)
by: Xiao, Chenghao, et al.
Published: (2025)
Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language Models
by: Leask, Patrick, et al.
Published: (2025)
by: Leask, Patrick, et al.
Published: (2025)
On Adversarial Robustness and Out-of-Distribution Robustness of Large Language Models
by: Yang, April, et al.
Published: (2024)
by: Yang, April, et al.
Published: (2024)
Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts
by: Hong, Hanhua, et al.
Published: (2025)
by: Hong, Hanhua, et al.
Published: (2025)
Evaluating Large Language Models for Generalization and Robustness via Data Compression
by: Li, Yucheng, et al.
Published: (2024)
by: Li, Yucheng, et al.
Published: (2024)
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation
by: James, Joseph, et al.
Published: (2026)
by: James, Joseph, et al.
Published: (2026)
Assessing Adversarial Robustness of Large Language Models: An Empirical Study
by: Yang, Zeyu, et al.
Published: (2024)
by: Yang, Zeyu, et al.
Published: (2024)
Robustness of Large Language Models Against Adversarial Attacks
by: Tao, Yiyi, et al.
Published: (2024)
by: Tao, Yiyi, et al.
Published: (2024)
Overview of the NLPCC 2024 Shared Task on Chinese Metaphor Generation
by: Qu, Xingwei, et al.
Published: (2024)
by: Qu, Xingwei, et al.
Published: (2024)
Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge
by: Li, Yizhi, et al.
Published: (2025)
by: Li, Yizhi, et al.
Published: (2025)
CAST: Corpus-Aware Self-similarity Enhanced Topic modelling
by: Ma, Yanan, et al.
Published: (2024)
by: Ma, Yanan, et al.
Published: (2024)
Defence Technology
Published: (2015)
Published: (2015)
BioMNER: A Dataset for Biomedical Method Entity Recognition
by: Tang, Chen, et al.
Published: (2024)
by: Tang, Chen, et al.
Published: (2024)
AttenCraft: Attention-guided Disentanglement of Multiple Concepts for Text-to-Image Customization
by: Shentu, Junjie, et al.
Published: (2024)
by: Shentu, Junjie, et al.
Published: (2024)
Textual Localization: Decomposing Multi-concept Images for Subject-Driven Text-to-Image Generation
by: Shentu, Junjie, et al.
Published: (2024)
by: Shentu, Junjie, et al.
Published: (2024)
Enhancing Language Agent Strategic Reasoning through Self-Play in Adversarial Games
by: Zhang, Yikai, et al.
Published: (2025)
by: Zhang, Yikai, et al.
Published: (2025)
PromptFix: Few-shot Backdoor Removal via Adversarial Prompt Tuning
by: Zhang, Tianrong, et al.
Published: (2024)
by: Zhang, Tianrong, et al.
Published: (2024)
Defensive Dual Masking for Robust Adversarial Defense
by: Yang, Wangli, et al.
Published: (2024)
by: Yang, Wangli, et al.
Published: (2024)
Nürnberg NLP at PsyDefDetect: Multi-Axis Voter Ensembles for Psychological Defence Mechanism Classification
by: Steigerwald, Philipp, et al.
Published: (2026)
by: Steigerwald, Philipp, et al.
Published: (2026)
Are All Prompt Components Value-Neutral? Understanding the Heterogeneous Adversarial Robustness of Dissected Prompt in Large Language Models
by: Zheng, Yujia, et al.
Published: (2025)
by: Zheng, Yujia, et al.
Published: (2025)
Natural Language Generation
by: van Miltenburg, Emiel, et al.
Published: (2025)
by: van Miltenburg, Emiel, et al.
Published: (2025)
Similar Items
-
Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence Benchmarks
by: Wang, Yang, et al.
Published: (2025) -
Audio Contrastive-based Fine-tuning: Decoupling Representation Learning and Classification
by: Wang, Yang, et al.
Published: (2023) -
RAR-b: Reasoning as Retrieval Benchmark
by: Xiao, Chenghao, et al.
Published: (2024) -
X-ray Made Simple: Lay Radiology Report Generation and Robust Evaluation
by: Zhao, Kun, et al.
Published: (2024) -
Pixel Sentence Representation Learning
by: Xiao, Chenghao, et al.
Published: (2024)