Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Si, Wai Man, Li, Mingjie, Backes, Michael, Zhang, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Excessive Reasoning Attack on Reasoning LLMs
by: Si, Wai Man, et al.
Published: (2025)
by: Si, Wai Man, et al.
Published: (2025)
A Systematic Study of Training-Free Methods for Trustworthy Large Language Models
by: Si, Wai Man, et al.
Published: (2026)
by: Si, Wai Man, et al.
Published: (2026)
Finding and Reactivating Post-Trained LLMs' Hidden Safety Mechanisms
by: Li, Mingjie, et al.
Published: (2026)
by: Li, Mingjie, et al.
Published: (2026)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation
by: Li, Mingjie, et al.
Published: (2025)
by: Li, Mingjie, et al.
Published: (2025)
VOCABTRIM: Vocabulary Pruning for Efficient Speculative Decoding in LLMs
by: Goel, Raghavv, et al.
Published: (2025)
by: Goel, Raghavv, et al.
Published: (2025)
ICLGuard: Controlling In-Context Learning Behavior for Applicability Authorization
by: Si, Wai Man, et al.
Published: (2024)
by: Si, Wai Man, et al.
Published: (2024)
On Pruning State-Space LLMs
by: Ghattas, Tamer, et al.
Published: (2025)
by: Ghattas, Tamer, et al.
Published: (2025)
On Importance of Pruning and Distillation for Efficient Low Resource NLP
by: Mirashi, Aishwarya, et al.
Published: (2024)
by: Mirashi, Aishwarya, et al.
Published: (2024)
Mixture Compressor for Mixture-of-Experts LLMs Gains More
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
by: Fu, Yao, et al.
Published: (2025)
by: Fu, Yao, et al.
Published: (2025)
DenoiseRotator: Enhance Pruning Robustness for LLMs via Importance Concentration
by: Gu, Tianteng, et al.
Published: (2025)
by: Gu, Tianteng, et al.
Published: (2025)
Random Masking Finds Winning Tickets for Parameter Efficient Fine-tuning
by: Xu, Jing, et al.
Published: (2024)
by: Xu, Jing, et al.
Published: (2024)
TrustLDM: Benchmarking Trustworthiness in Language Diffusion Models
by: Mo, Yichuan, et al.
Published: (2026)
by: Mo, Yichuan, et al.
Published: (2026)
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
by: Mendu, Sai Krishna, et al.
Published: (2025)
by: Mendu, Sai Krishna, et al.
Published: (2025)
Evaluating Defences against Unsafe Feedback in RLHF
by: Rosati, Domenic, et al.
Published: (2024)
by: Rosati, Domenic, et al.
Published: (2024)
Early Transformers: A study on Efficient Training of Transformer Models through Early-Bird Lottery Tickets
by: Cheekati, Shravan
Published: (2024)
by: Cheekati, Shravan
Published: (2024)
Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
by: Kolawole, Steven, et al.
Published: (2024)
by: Kolawole, Steven, et al.
Published: (2024)
Learning and Forgetting Unsafe Examples in Large Language Models
by: Zhao, Jiachen, et al.
Published: (2023)
by: Zhao, Jiachen, et al.
Published: (2023)
Pruning Strategies for Backdoor Defense in LLMs
by: Chapagain, Santosh, et al.
Published: (2025)
by: Chapagain, Santosh, et al.
Published: (2025)
A Simple and Effective Pruning Approach for Large Language Models
by: Sun, Mingjie, et al.
Published: (2023)
by: Sun, Mingjie, et al.
Published: (2023)
Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-Probing
by: Le, Qi, et al.
Published: (2025)
by: Le, Qi, et al.
Published: (2025)
2SSP: A Two-Stage Framework for Structured Pruning of LLMs
by: Sandri, Fabrizio, et al.
Published: (2025)
by: Sandri, Fabrizio, et al.
Published: (2025)
Safer Policy Compliance with Dynamic Epistemic Fallback
by: Imperial, Joseph Marvin, et al.
Published: (2026)
by: Imperial, Joseph Marvin, et al.
Published: (2026)
Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts
by: Ye, Jiayuan, et al.
Published: (2026)
by: Ye, Jiayuan, et al.
Published: (2026)
EfficientXpert: Efficient Domain Adaptation for Large Language Models via Propagation-Aware Pruning
by: Zhao, Songlin, et al.
Published: (2025)
by: Zhao, Songlin, et al.
Published: (2025)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
SOS! Soft Prompt Attack Against Open-Source Large Language Models
by: Yang, Ziqing, et al.
Published: (2024)
by: Yang, Ziqing, et al.
Published: (2024)
Towards Robust and Parameter-Efficient Knowledge Unlearning for LLMs
by: Cha, Sungmin, et al.
Published: (2024)
by: Cha, Sungmin, et al.
Published: (2024)
FaultProfIT: Hierarchical Fault Profiling of Incident Tickets in Large-scale Cloud Systems
by: Huang, Junjie, et al.
Published: (2024)
by: Huang, Junjie, et al.
Published: (2024)
LLMGuard: Guarding Against Unsafe LLM Behavior
by: Goyal, Shubh, et al.
Published: (2024)
by: Goyal, Shubh, et al.
Published: (2024)
Beyond Size: How Gradients Shape Pruning Decisions in Large Language Models
by: Das, Rocktim Jyoti, et al.
Published: (2023)
by: Das, Rocktim Jyoti, et al.
Published: (2023)
Two-stage LLM Fine-tuning with Less Specialization and More Generalization
by: Wang, Yihan, et al.
Published: (2022)
by: Wang, Yihan, et al.
Published: (2022)
ROSE: Reordered SparseGPT for More Accurate One-Shot Large Language Models Pruning
by: Su, Mingluo, et al.
Published: (2026)
by: Su, Mingluo, et al.
Published: (2026)
Adaptive LoRA Merge with Parameter Pruning for Low-Resource Generation
by: Miyano, Ryota, et al.
Published: (2025)
by: Miyano, Ryota, et al.
Published: (2025)
On Importance of Layer Pruning for Smaller BERT Models and Low Resource Languages
by: Shirke, Mayur, et al.
Published: (2025)
by: Shirke, Mayur, et al.
Published: (2025)
MoreauPruner: Robust Pruning of Large Language Models against Weight Perturbations
by: Wang, Zixiao, et al.
Published: (2024)
by: Wang, Zixiao, et al.
Published: (2024)
KS-Lottery: Finding Certified Lottery Tickets for Multilingual Language Models
by: Yuan, Fei, et al.
Published: (2024)
by: Yuan, Fei, et al.
Published: (2024)
Less is More: Extreme Gradient Boost Rank-1 Adaption for Efficient Finetuning of LLMs
by: Zhang, Yifei, et al.
Published: (2024)
by: Zhang, Yifei, et al.
Published: (2024)
Why Are Web AI Agents More Vulnerable Than Standalone LLMs? A Security Analysis
by: Chiang, Jeffrey Yang Fan, et al.
Published: (2025)
by: Chiang, Jeffrey Yang Fan, et al.
Published: (2025)
Similar Items
-
Excessive Reasoning Attack on Reasoning LLMs
by: Si, Wai Man, et al.
Published: (2025) -
A Systematic Study of Training-Free Methods for Trustworthy Large Language Models
by: Si, Wai Man, et al.
Published: (2026) -
Finding and Reactivating Post-Trained LLMs' Hidden Safety Mechanisms
by: Li, Mingjie, et al.
Published: (2026) -
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
by: Jiang, Yukun, et al.
Published: (2026) -
SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation
by: Li, Mingjie, et al.
Published: (2025)