Investigating the Impact of Quantization Methods on the Safety and Reliability of Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kharinaev, Artyom, Moskvoretskii, Viktor, Shvetsov, Egor, Studenikina, Kseniia, Mikhail, Bykov, Burnaev, Evgeny |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EBES: Easy Benchmarking for Event Sequences
von: Osin, Dmitry, et al.
Veröffentlicht: (2024)
von: Osin, Dmitry, et al.
Veröffentlicht: (2024)
Investigating the Impact of Quantization on Adversarial Robustness
von: Li, Qun, et al.
Veröffentlicht: (2024)
von: Li, Qun, et al.
Veröffentlicht: (2024)
GIFT-SW: Gaussian noise Injected Fine-Tuning of Salient Weights for LLMs
von: Zhelnin, Maxim, et al.
Veröffentlicht: (2024)
von: Zhelnin, Maxim, et al.
Veröffentlicht: (2024)
Q-MLLM: Vector Quantization for Robust Multimodal Large Language Model Security
von: Zhao, Wei, et al.
Veröffentlicht: (2025)
von: Zhao, Wei, et al.
Veröffentlicht: (2025)
A Cross-Language Investigation into Jailbreak Attacks in Large Language Models
von: Li, Jie, et al.
Veröffentlicht: (2024)
von: Li, Jie, et al.
Veröffentlicht: (2024)
Protecting Private Code in IDE Autocomplete using Differential Privacy
von: Grigorenko, Evgeny, et al.
Veröffentlicht: (2026)
von: Grigorenko, Evgeny, et al.
Veröffentlicht: (2026)
Safety Layers in Aligned Large Language Models: The Key to LLM Security
von: Li, Shen, et al.
Veröffentlicht: (2024)
von: Li, Shen, et al.
Veröffentlicht: (2024)
Robustifying Safety-Aligned Large Language Models through Clean Data Curation
von: Liu, Xiaoqun, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoqun, et al.
Veröffentlicht: (2024)
A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense
von: Zhai, Keke
Veröffentlicht: (2024)
von: Zhai, Keke
Veröffentlicht: (2024)
Exploring the Potential of Large Language Models for Improving Digital Forensic Investigation Efficiency
von: Wickramasekara, Akila, et al.
Veröffentlicht: (2024)
von: Wickramasekara, Akila, et al.
Veröffentlicht: (2024)
Depth Charge: Jailbreak Large Language Models from Deep Safety Attention Heads
von: Wu, Jinman, et al.
Veröffentlicht: (2026)
von: Wu, Jinman, et al.
Veröffentlicht: (2026)
SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models
von: Chen, Yulong, et al.
Veröffentlicht: (2025)
von: Chen, Yulong, et al.
Veröffentlicht: (2025)
AISA: Awakening Intrinsic Safety Awareness in Large Language Models against Jailbreak Attacks
von: Song, Weiming, et al.
Veröffentlicht: (2026)
von: Song, Weiming, et al.
Veröffentlicht: (2026)
USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models
von: Zheng, Baolin, et al.
Veröffentlicht: (2025)
von: Zheng, Baolin, et al.
Veröffentlicht: (2025)
Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models
von: Wang, Xunguang, et al.
Veröffentlicht: (2026)
von: Wang, Xunguang, et al.
Veröffentlicht: (2026)
Differentially Private and Communication Efficient Large Language Model Split Inference via Stochastic Quantization and Soft Prompt
von: Gu, Yujie, et al.
Veröffentlicht: (2026)
von: Gu, Yujie, et al.
Veröffentlicht: (2026)
On Jailbreaking Quantized Language Models Through Fault Injection Attacks
von: Zahran, Noureldin, et al.
Veröffentlicht: (2025)
von: Zahran, Noureldin, et al.
Veröffentlicht: (2025)
Can Transformer Memory Be Corrupted? Investigating Cache-Side Vulnerabilities in Large Language Models
von: Hossain, Elias, et al.
Veröffentlicht: (2025)
von: Hossain, Elias, et al.
Veröffentlicht: (2025)
Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
FAA Framework: A Large Language Model-Based Approach for Credit Card Fraud Investigations
von: Shuster, Shaun, et al.
Veröffentlicht: (2025)
von: Shuster, Shaun, et al.
Veröffentlicht: (2025)
Enhancing Reverse Engineering: Investigating and Benchmarking Large Language Models for Vulnerability Analysis in Decompiled Binaries
von: Manuel, Dylan, et al.
Veröffentlicht: (2024)
von: Manuel, Dylan, et al.
Veröffentlicht: (2024)
On the Reliability and Stability of Selective Methods in Malware Classification Tasks
von: Herzog, Alexander, et al.
Veröffentlicht: (2025)
von: Herzog, Alexander, et al.
Veröffentlicht: (2025)
Structured Visual Narratives Undermine Safety Alignment in Multimodal Large Language Models
von: Tan, Rui Yang, et al.
Veröffentlicht: (2026)
von: Tan, Rui Yang, et al.
Veröffentlicht: (2026)
Understanding the Effects of Safety Unalignment on Large Language Models
von: Halloran, John T.
Veröffentlicht: (2026)
von: Halloran, John T.
Veröffentlicht: (2026)
Internal Safety Collapse in Frontier Large Language Models
von: Wu, Yutao, et al.
Veröffentlicht: (2026)
von: Wu, Yutao, et al.
Veröffentlicht: (2026)
Quantized Delta Weight Is Safety Keeper
von: Liu, Yule, et al.
Veröffentlicht: (2024)
von: Liu, Yule, et al.
Veröffentlicht: (2024)
DrLLM: Prompt-Enhanced Distributed Denial-of-Service Resistance Method with Large Language Models
von: Yin, Zhenyu, et al.
Veröffentlicht: (2024)
von: Yin, Zhenyu, et al.
Veröffentlicht: (2024)
Finetuning Large Language Models for Vulnerability Detection
von: Shestov, Alexey, et al.
Veröffentlicht: (2024)
von: Shestov, Alexey, et al.
Veröffentlicht: (2024)
Evaluating the Reliability and Fidelity of Automated Judgment Systems of Large Language Models
von: Biskupski, Tom, et al.
Veröffentlicht: (2026)
von: Biskupski, Tom, et al.
Veröffentlicht: (2026)
Can Small Language Models Reliably Resist Jailbreak Attacks? A Comprehensive Evaluation
von: Zhang, Wenhui, et al.
Veröffentlicht: (2025)
von: Zhang, Wenhui, et al.
Veröffentlicht: (2025)
SGuard-v1: Safety Guardrail for Large Language Models
von: Lee, JoonHo, et al.
Veröffentlicht: (2025)
von: Lee, JoonHo, et al.
Veröffentlicht: (2025)
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications
von: David, Isaac, et al.
Veröffentlicht: (2026)
von: David, Isaac, et al.
Veröffentlicht: (2026)
State-Dependent Safety Failures in Multi-Turn Language Model Interaction
von: Li, Pengcheng, et al.
Veröffentlicht: (2026)
von: Li, Pengcheng, et al.
Veröffentlicht: (2026)
(Security) Assertions by Large Language Models
von: Kande, Rahul, et al.
Veröffentlicht: (2023)
von: Kande, Rahul, et al.
Veröffentlicht: (2023)
LLMmap: Fingerprinting For Large Language Models
von: Pasquini, Dario, et al.
Veröffentlicht: (2024)
von: Pasquini, Dario, et al.
Veröffentlicht: (2024)
Backdooring Bias in Large Language Models
von: Das, Anudeep, et al.
Veröffentlicht: (2026)
von: Das, Anudeep, et al.
Veröffentlicht: (2026)
ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models
von: Owiredu-Ashley, Harry
Veröffentlicht: (2026)
von: Owiredu-Ashley, Harry
Veröffentlicht: (2026)
BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models
von: Zeng, Yi, et al.
Veröffentlicht: (2024)
von: Zeng, Yi, et al.
Veröffentlicht: (2024)
Federated Learning-Based Data Collaboration Method for Enhancing Edge Cloud AI System Security Using Large Language Models
von: Luo, Huaiying, et al.
Veröffentlicht: (2025)
von: Luo, Huaiying, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
EBES: Easy Benchmarking for Event Sequences
von: Osin, Dmitry, et al.
Veröffentlicht: (2024) -
Investigating the Impact of Quantization on Adversarial Robustness
von: Li, Qun, et al.
Veröffentlicht: (2024) -
GIFT-SW: Gaussian noise Injected Fine-Tuning of Salient Weights for LLMs
von: Zhelnin, Maxim, et al.
Veröffentlicht: (2024) -
Q-MLLM: Vector Quantization for Robust Multimodal Large Language Model Security
von: Zhao, Wei, et al.
Veröffentlicht: (2025) -
A Cross-Language Investigation into Jailbreak Attacks in Large Language Models
von: Li, Jie, et al.
Veröffentlicht: (2024)