Pruning Strategies for Backdoor Defense in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Chapagain, Santosh, Hamdi, Shah Muhammad, Boubrahimi, Soukaina Filali |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Advancing Hate Speech Detection with Transformers: Insights from the MetaHate
by: Chapagain, Santosh, et al.
Published: (2025)
by: Chapagain, Santosh, et al.
Published: (2025)
SolarGPT-QA: A Domain-Adaptive Large Language Model for Educational Question Answering in Space Weather and Heliophysics
by: Chapagain, Santosh, et al.
Published: (2026)
by: Chapagain, Santosh, et al.
Published: (2026)
Advancing Minority Stress Detection with Transformers: Insights from the Social Media Datasets
by: Chapagain, Santosh, et al.
Published: (2025)
by: Chapagain, Santosh, et al.
Published: (2025)
Motif-guided Time Series Counterfactual Explanations
by: Li, Peiyu, et al.
Published: (2022)
by: Li, Peiyu, et al.
Published: (2022)
Global Cross-Time Attention Fusion for Enhanced Solar Flare Prediction from Multivariate Time Series
by: Vural, Onur, et al.
Published: (2025)
by: Vural, Onur, et al.
Published: (2025)
EXCON: Extreme Instance-based Contrastive Representation Learning of Severely Imbalanced Multivariate Time Series for Solar Flare Prediction
by: Vural, Onur, et al.
Published: (2024)
by: Vural, Onur, et al.
Published: (2024)
TIMED: Adversarial and Autoregressive Refinement of Diffusion-Based Time Series Generation
by: EskandariNasab, MohammadReza, et al.
Published: (2025)
by: EskandariNasab, MohammadReza, et al.
Published: (2025)
Contrastive Representation Learning for Predicting Solar Flares from Extremely Imbalanced Multivariate Time Series Data
by: Vural, Onur, et al.
Published: (2024)
by: Vural, Onur, et al.
Published: (2024)
AVATAR: Adversarial Autoencoders with Autoregressive Refinement for Time Series Generation
by: EskandariNasab, MohammadReza, et al.
Published: (2025)
by: EskandariNasab, MohammadReza, et al.
Published: (2025)
SeriesGAN: Time Series Generation via Adversarial and Autoregressive Learning
by: EskandariNasab, MohammadReza, et al.
Published: (2024)
by: EskandariNasab, MohammadReza, et al.
Published: (2024)
ChronoGAN: Supervised and Embedded Generative Adversarial Networks for Time Series Generation
by: EskandariNasab, MohammadReza, et al.
Published: (2024)
by: EskandariNasab, MohammadReza, et al.
Published: (2024)
M-CELS: Counterfactual Explanation for Multivariate Time Series Data Guided by Learned Saliency Maps
by: Li, Peiyu, et al.
Published: (2024)
by: Li, Peiyu, et al.
Published: (2024)
Enhancing Multivariate Time Series-based Solar Flare Prediction with Multifaceted Preprocessing and Contrastive Learning
by: EskandariNasab, MohammadReza, et al.
Published: (2024)
by: EskandariNasab, MohammadReza, et al.
Published: (2024)
Info-CELS: Informative Saliency Map Guided Counterfactual Explanation
by: Li, Peiyu, et al.
Published: (2024)
by: Li, Peiyu, et al.
Published: (2024)
Forest Proximities for Time Series
by: Shaw, Ben, et al.
Published: (2024)
by: Shaw, Ben, et al.
Published: (2024)
Source Code in Support of Solar Flare Prediction through Time Series Data Augmentation
by: Li, Peiyu, et al.
Published: (2025)
by: Li, Peiyu, et al.
Published: (2025)
Confidence-Credibility Aware Weighted Ensembles of Small LLMs Outperform Large LLMs in Emotion Detection
by: Elgabry, Menna, et al.
Published: (2025)
by: Elgabry, Menna, et al.
Published: (2025)
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
by: Fu, Yao, et al.
Published: (2025)
by: Fu, Yao, et al.
Published: (2025)
On Pruning State-Space LLMs
by: Ghattas, Tamer, et al.
Published: (2025)
by: Ghattas, Tamer, et al.
Published: (2025)
Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
by: Kolawole, Steven, et al.
Published: (2024)
by: Kolawole, Steven, et al.
Published: (2024)
Data-centric NLP Backdoor Defense from the Lens of Memorization
by: Wang, Zhenting, et al.
Published: (2024)
by: Wang, Zhenting, et al.
Published: (2024)
Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language Models
by: Das, Anindya Sundar, et al.
Published: (2025)
by: Das, Anindya Sundar, et al.
Published: (2025)
Pruning as a Defense: Reducing Memorization in Large Language Models
by: Gupta, Mansi, et al.
Published: (2025)
by: Gupta, Mansi, et al.
Published: (2025)
VOCABTRIM: Vocabulary Pruning for Efficient Speculative Decoding in LLMs
by: Goel, Raghavv, et al.
Published: (2025)
by: Goel, Raghavv, et al.
Published: (2025)
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
by: Liu, Qin, et al.
Published: (2023)
by: Liu, Qin, et al.
Published: (2023)
DenoiseRotator: Enhance Pruning Robustness for LLMs via Importance Concentration
by: Gu, Tianteng, et al.
Published: (2025)
by: Gu, Tianteng, et al.
Published: (2025)
Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-Probing
by: Le, Qi, et al.
Published: (2025)
by: Le, Qi, et al.
Published: (2025)
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
by: Sivapiromrat, Sanhanat, et al.
Published: (2025)
by: Sivapiromrat, Sanhanat, et al.
Published: (2025)
Dynamic Vocabulary Pruning in Early-Exit LLMs
by: Vincenti, Jort, et al.
Published: (2024)
by: Vincenti, Jort, et al.
Published: (2024)
Adaptive Multi-Expert Reasoning via Difficulty-Aware Routing and Uncertainty-Guided Aggregation
by: Ehab, Mohamed, et al.
Published: (2026)
by: Ehab, Mohamed, et al.
Published: (2026)
Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
by: Price, Sara, et al.
Published: (2024)
by: Price, Sara, et al.
Published: (2024)
Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs
by: Si, Wai Man, et al.
Published: (2026)
by: Si, Wai Man, et al.
Published: (2026)
Two-Stage Regularization-Based Structured Pruning for LLMs
by: Feng, Mingkuan, et al.
Published: (2025)
by: Feng, Mingkuan, et al.
Published: (2025)
UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
by: Lin, Huawei, et al.
Published: (2025)
by: Lin, Huawei, et al.
Published: (2025)
REAM: Merging Improves Pruning of Experts in LLMs
by: Jha, Saurav, et al.
Published: (2026)
by: Jha, Saurav, et al.
Published: (2026)
STUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning
by: Lee, Jaeseong, et al.
Published: (2024)
by: Lee, Jaeseong, et al.
Published: (2024)
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
by: Brown, Hannah, et al.
Published: (2024)
by: Brown, Hannah, et al.
Published: (2024)
Breaking Down Bias: On The Limits of Generalizable Pruning Strategies
by: Ma, Sibo, et al.
Published: (2025)
by: Ma, Sibo, et al.
Published: (2025)
CAMO: A Class-Aware Minority-Optimized Ensemble for Robust Language Model Evaluation on Imbalanced Data
by: Ehab, Mohamed, et al.
Published: (2026)
by: Ehab, Mohamed, et al.
Published: (2026)
CMHL: Contrastive Multi-Head Learning for Emotionally Consistent Text Classification
by: Elgabry, Menna, et al.
Published: (2026)
by: Elgabry, Menna, et al.
Published: (2026)
Similar Items
-
Advancing Hate Speech Detection with Transformers: Insights from the MetaHate
by: Chapagain, Santosh, et al.
Published: (2025) -
SolarGPT-QA: A Domain-Adaptive Large Language Model for Educational Question Answering in Space Weather and Heliophysics
by: Chapagain, Santosh, et al.
Published: (2026) -
Advancing Minority Stress Detection with Transformers: Insights from the Social Media Datasets
by: Chapagain, Santosh, et al.
Published: (2025) -
Motif-guided Time Series Counterfactual Explanations
by: Li, Peiyu, et al.
Published: (2022) -
Global Cross-Time Attention Fusion for Enhanced Solar Flare Prediction from Multivariate Time Series
by: Vural, Onur, et al.
Published: (2025)