Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Si, Wai Man, Li, Mingjie, Backes, Michael, Zhang, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Excessive Reasoning Attack on Reasoning LLMs
von: Si, Wai Man, et al.
Veröffentlicht: (2025)
von: Si, Wai Man, et al.
Veröffentlicht: (2025)
A Systematic Study of Training-Free Methods for Trustworthy Large Language Models
von: Si, Wai Man, et al.
Veröffentlicht: (2026)
von: Si, Wai Man, et al.
Veröffentlicht: (2026)
Finding and Reactivating Post-Trained LLMs' Hidden Safety Mechanisms
von: Li, Mingjie, et al.
Veröffentlicht: (2026)
von: Li, Mingjie, et al.
Veröffentlicht: (2026)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
von: Jiang, Yukun, et al.
Veröffentlicht: (2026)
von: Jiang, Yukun, et al.
Veröffentlicht: (2026)
SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation
von: Li, Mingjie, et al.
Veröffentlicht: (2025)
von: Li, Mingjie, et al.
Veröffentlicht: (2025)
VOCABTRIM: Vocabulary Pruning for Efficient Speculative Decoding in LLMs
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
ICLGuard: Controlling In-Context Learning Behavior for Applicability Authorization
von: Si, Wai Man, et al.
Veröffentlicht: (2024)
von: Si, Wai Man, et al.
Veröffentlicht: (2024)
On Pruning State-Space LLMs
von: Ghattas, Tamer, et al.
Veröffentlicht: (2025)
von: Ghattas, Tamer, et al.
Veröffentlicht: (2025)
On Importance of Pruning and Distillation for Efficient Low Resource NLP
von: Mirashi, Aishwarya, et al.
Veröffentlicht: (2024)
von: Mirashi, Aishwarya, et al.
Veröffentlicht: (2024)
Mixture Compressor for Mixture-of-Experts LLMs Gains More
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
von: Fu, Yao, et al.
Veröffentlicht: (2025)
von: Fu, Yao, et al.
Veröffentlicht: (2025)
DenoiseRotator: Enhance Pruning Robustness for LLMs via Importance Concentration
von: Gu, Tianteng, et al.
Veröffentlicht: (2025)
von: Gu, Tianteng, et al.
Veröffentlicht: (2025)
Random Masking Finds Winning Tickets for Parameter Efficient Fine-tuning
von: Xu, Jing, et al.
Veröffentlicht: (2024)
von: Xu, Jing, et al.
Veröffentlicht: (2024)
TrustLDM: Benchmarking Trustworthiness in Language Diffusion Models
von: Mo, Yichuan, et al.
Veröffentlicht: (2026)
von: Mo, Yichuan, et al.
Veröffentlicht: (2026)
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
von: Mendu, Sai Krishna, et al.
Veröffentlicht: (2025)
von: Mendu, Sai Krishna, et al.
Veröffentlicht: (2025)
Evaluating Defences against Unsafe Feedback in RLHF
von: Rosati, Domenic, et al.
Veröffentlicht: (2024)
von: Rosati, Domenic, et al.
Veröffentlicht: (2024)
Early Transformers: A study on Efficient Training of Transformer Models through Early-Bird Lottery Tickets
von: Cheekati, Shravan
Veröffentlicht: (2024)
von: Cheekati, Shravan
Veröffentlicht: (2024)
Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
von: Kolawole, Steven, et al.
Veröffentlicht: (2024)
von: Kolawole, Steven, et al.
Veröffentlicht: (2024)
Learning and Forgetting Unsafe Examples in Large Language Models
von: Zhao, Jiachen, et al.
Veröffentlicht: (2023)
von: Zhao, Jiachen, et al.
Veröffentlicht: (2023)
Pruning Strategies for Backdoor Defense in LLMs
von: Chapagain, Santosh, et al.
Veröffentlicht: (2025)
von: Chapagain, Santosh, et al.
Veröffentlicht: (2025)
A Simple and Effective Pruning Approach for Large Language Models
von: Sun, Mingjie, et al.
Veröffentlicht: (2023)
von: Sun, Mingjie, et al.
Veröffentlicht: (2023)
Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-Probing
von: Le, Qi, et al.
Veröffentlicht: (2025)
von: Le, Qi, et al.
Veröffentlicht: (2025)
2SSP: A Two-Stage Framework for Structured Pruning of LLMs
von: Sandri, Fabrizio, et al.
Veröffentlicht: (2025)
von: Sandri, Fabrizio, et al.
Veröffentlicht: (2025)
Safer Policy Compliance with Dynamic Epistemic Fallback
von: Imperial, Joseph Marvin, et al.
Veröffentlicht: (2026)
von: Imperial, Joseph Marvin, et al.
Veröffentlicht: (2026)
Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts
von: Ye, Jiayuan, et al.
Veröffentlicht: (2026)
von: Ye, Jiayuan, et al.
Veröffentlicht: (2026)
EfficientXpert: Efficient Domain Adaptation for Large Language Models via Propagation-Aware Pruning
von: Zhao, Songlin, et al.
Veröffentlicht: (2025)
von: Zhao, Songlin, et al.
Veröffentlicht: (2025)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
SOS! Soft Prompt Attack Against Open-Source Large Language Models
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
Towards Robust and Parameter-Efficient Knowledge Unlearning for LLMs
von: Cha, Sungmin, et al.
Veröffentlicht: (2024)
von: Cha, Sungmin, et al.
Veröffentlicht: (2024)
FaultProfIT: Hierarchical Fault Profiling of Incident Tickets in Large-scale Cloud Systems
von: Huang, Junjie, et al.
Veröffentlicht: (2024)
von: Huang, Junjie, et al.
Veröffentlicht: (2024)
LLMGuard: Guarding Against Unsafe LLM Behavior
von: Goyal, Shubh, et al.
Veröffentlicht: (2024)
von: Goyal, Shubh, et al.
Veröffentlicht: (2024)
Beyond Size: How Gradients Shape Pruning Decisions in Large Language Models
von: Das, Rocktim Jyoti, et al.
Veröffentlicht: (2023)
von: Das, Rocktim Jyoti, et al.
Veröffentlicht: (2023)
Two-stage LLM Fine-tuning with Less Specialization and More Generalization
von: Wang, Yihan, et al.
Veröffentlicht: (2022)
von: Wang, Yihan, et al.
Veröffentlicht: (2022)
ROSE: Reordered SparseGPT for More Accurate One-Shot Large Language Models Pruning
von: Su, Mingluo, et al.
Veröffentlicht: (2026)
von: Su, Mingluo, et al.
Veröffentlicht: (2026)
Adaptive LoRA Merge with Parameter Pruning for Low-Resource Generation
von: Miyano, Ryota, et al.
Veröffentlicht: (2025)
von: Miyano, Ryota, et al.
Veröffentlicht: (2025)
On Importance of Layer Pruning for Smaller BERT Models and Low Resource Languages
von: Shirke, Mayur, et al.
Veröffentlicht: (2025)
von: Shirke, Mayur, et al.
Veröffentlicht: (2025)
MoreauPruner: Robust Pruning of Large Language Models against Weight Perturbations
von: Wang, Zixiao, et al.
Veröffentlicht: (2024)
von: Wang, Zixiao, et al.
Veröffentlicht: (2024)
KS-Lottery: Finding Certified Lottery Tickets for Multilingual Language Models
von: Yuan, Fei, et al.
Veröffentlicht: (2024)
von: Yuan, Fei, et al.
Veröffentlicht: (2024)
Less is More: Extreme Gradient Boost Rank-1 Adaption for Efficient Finetuning of LLMs
von: Zhang, Yifei, et al.
Veröffentlicht: (2024)
von: Zhang, Yifei, et al.
Veröffentlicht: (2024)
Why Are Web AI Agents More Vulnerable Than Standalone LLMs? A Security Analysis
von: Chiang, Jeffrey Yang Fan, et al.
Veröffentlicht: (2025)
von: Chiang, Jeffrey Yang Fan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Excessive Reasoning Attack on Reasoning LLMs
von: Si, Wai Man, et al.
Veröffentlicht: (2025) -
A Systematic Study of Training-Free Methods for Trustworthy Large Language Models
von: Si, Wai Man, et al.
Veröffentlicht: (2026) -
Finding and Reactivating Post-Trained LLMs' Hidden Safety Mechanisms
von: Li, Mingjie, et al.
Veröffentlicht: (2026) -
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
von: Jiang, Yukun, et al.
Veröffentlicht: (2026) -
SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation
von: Li, Mingjie, et al.
Veröffentlicht: (2025)