Locking Down the Finetuned LLMs Safety
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Minjun, Yang, Linyi, Wei, Yifan, Zhang, Ningyu, Zhang, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Personality Alignment of Large Language Models
by: Zhu, Minjun, et al.
Published: (2024)
by: Zhu, Minjun, et al.
Published: (2024)
DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process
by: Zhu, Minjun, et al.
Published: (2025)
by: Zhu, Minjun, et al.
Published: (2025)
MCEval: A Dynamic Framework for Fair Multilingual Cultural Evaluation of LLMs
by: Huang, Shulin, et al.
Published: (2025)
by: Huang, Shulin, et al.
Published: (2025)
CycleResearcher: Improving Automated Research via Automated Review
by: Weng, Yixuan, et al.
Published: (2024)
by: Weng, Yixuan, et al.
Published: (2024)
AI Scientists Fail Without Strong Implementation Capability
by: Zhu, Minjun, et al.
Published: (2025)
by: Zhu, Minjun, et al.
Published: (2025)
SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts
by: Xin, Yuan, et al.
Published: (2026)
by: Xin, Yuan, et al.
Published: (2026)
How Likely Do LLMs with CoT Mimic Human Reasoning?
by: Bao, Guangsheng, et al.
Published: (2024)
by: Bao, Guangsheng, et al.
Published: (2024)
Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined Values
by: Zhang, Hongbo, et al.
Published: (2025)
by: Zhang, Hongbo, et al.
Published: (2025)
Identity Lock: Locking API Fine-tuned LLMs With Identity-based Wake Words
by: Su, Hongyu, et al.
Published: (2025)
by: Su, Hongyu, et al.
Published: (2025)
Understanding the Effects of Domain Finetuning on LLMs
by: Tanwar, Eshaan, et al.
Published: (2025)
by: Tanwar, Eshaan, et al.
Published: (2025)
Safety-Aligned Weights Are Not Enough: Refusal-Teacher-Guided Finetuning Enhances Safety and Downstream Performance under Harmful Finetuning Attacks
by: Ham, Seokil, et al.
Published: (2025)
by: Ham, Seokil, et al.
Published: (2025)
Constrain Alignment with Sparse Autoencoders
by: Yin, Qingyu, et al.
Published: (2024)
by: Yin, Qingyu, et al.
Published: (2024)
T-Detect: Tail-Aware Statistical Normalization for Robust Detection of Adversarial Machine-Generated Text
by: West, Alva, et al.
Published: (2025)
by: West, Alva, et al.
Published: (2025)
Finetuning LLMs for Comparative Assessment Tasks
by: Raina, Vatsal, et al.
Published: (2024)
by: Raina, Vatsal, et al.
Published: (2024)
Cofca: A Step-Wise Counterfactual Multi-hop QA benchmark
by: Wu, Jian, et al.
Published: (2024)
by: Wu, Jian, et al.
Published: (2024)
Unveiling Over-Memorization in Finetuning LLMs for Reasoning Tasks
by: Ruan, Zhiwen, et al.
Published: (2025)
by: Ruan, Zhiwen, et al.
Published: (2025)
Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature
by: Bao, Guangsheng, et al.
Published: (2023)
by: Bao, Guangsheng, et al.
Published: (2023)
How LoRA Remembers? A Parametric Memory Law for LLM Finetuning
by: Xu, Ziwen, et al.
Published: (2026)
by: Xu, Ziwen, et al.
Published: (2026)
Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
by: Xin, Yuan, et al.
Published: (2025)
by: Xin, Yuan, et al.
Published: (2025)
Synthesizing Privacy-Preserving Text Data via Finetuning without Finetuning Billion-Scale LLMs
by: Tan, Bowen, et al.
Published: (2025)
by: Tan, Bowen, et al.
Published: (2025)
Improving Task Diversity in Label Efficient Supervised Finetuning of LLMs
by: Arabelly, Abhinav, et al.
Published: (2025)
by: Arabelly, Abhinav, et al.
Published: (2025)
ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails
by: Wang, Yan, et al.
Published: (2026)
by: Wang, Yan, et al.
Published: (2026)
Accurate LoRA-Finetuning Quantization of LLMs via Information Retention
by: Qin, Haotong, et al.
Published: (2024)
by: Qin, Haotong, et al.
Published: (2024)
Evaluation of Finetuned LLMs in AMR Parsing
by: Ho, Shu Han
Published: (2025)
by: Ho, Shu Han
Published: (2025)
A Rationale-centric Counterfactual Data Augmentation Method for Cross-Document Event Coreference Resolution
by: Ding, Bowen, et al.
Published: (2024)
by: Ding, Bowen, et al.
Published: (2024)
AI-Generated Text is Non-Stationary: Detection via Temporal Tomography
by: West, Alva, et al.
Published: (2025)
by: West, Alva, et al.
Published: (2025)
LANDeRMT: Detecting and Routing Language-Aware Neurons for Selectively Finetuning LLMs to Machine Translation
by: Zhu, Shaolin, et al.
Published: (2024)
by: Zhu, Shaolin, et al.
Published: (2024)
An Empirical Analysis of Uncertainty in Large Language Model Evaluations
by: Xie, Qiujie, et al.
Published: (2025)
by: Xie, Qiujie, et al.
Published: (2025)
Less is More: Extreme Gradient Boost Rank-1 Adaption for Efficient Finetuning of LLMs
by: Zhang, Yifei, et al.
Published: (2024)
by: Zhang, Yifei, et al.
Published: (2024)
SASFT: Sparse Autoencoder-guided Supervised Finetuning to Mitigate Unexpected Code-Switching in LLMs
by: Deng, Boyi, et al.
Published: (2025)
by: Deng, Boyi, et al.
Published: (2025)
Abduct, Act, Predict: Scaffolding Causal Inference for Automated Failure Attribution in Multi-Agent Systems
by: West, Alva, et al.
Published: (2025)
by: West, Alva, et al.
Published: (2025)
MemLLM: Finetuning LLMs to Use An Explicit Read-Write Memory
by: Modarressi, Ali, et al.
Published: (2024)
by: Modarressi, Ali, et al.
Published: (2024)
Finetuning LLMs for EvaCun 2025 token prediction shared task
by: Jon, Josef, et al.
Published: (2025)
by: Jon, Josef, et al.
Published: (2025)
LLMs and Finetuning: Benchmarking cross-domain performance for hate speech detection
by: Nasir, Ahmad, et al.
Published: (2023)
by: Nasir, Ahmad, et al.
Published: (2023)
When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method
by: Zhang, Biao, et al.
Published: (2024)
by: Zhang, Biao, et al.
Published: (2024)
CAP: Data Contamination Detection via Consistency Amplification
by: Zhao, Yi, et al.
Published: (2024)
by: Zhao, Yi, et al.
Published: (2024)
Lighting Up or Dimming Down? Exploring Dark Patterns of LLMs in Co-Creativity
by: Li, Zhu, et al.
Published: (2026)
by: Li, Zhu, et al.
Published: (2026)
AutoFigure: Generating and Refining Publication-Ready Scientific Illustrations
by: Zhu, Minjun, et al.
Published: (2026)
by: Zhu, Minjun, et al.
Published: (2026)
MAGE: Machine-generated Text Detection in the Wild
by: Li, Yafu, et al.
Published: (2023)
by: Li, Yafu, et al.
Published: (2023)
Ask, Answer, and Detect: Role-Playing LLMs for Personality Detection with Question-Conditioned Mixture-of-Experts
by: Lyu, Yifan, et al.
Published: (2025)
by: Lyu, Yifan, et al.
Published: (2025)
Similar Items
-
Personality Alignment of Large Language Models
by: Zhu, Minjun, et al.
Published: (2024) -
DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process
by: Zhu, Minjun, et al.
Published: (2025) -
MCEval: A Dynamic Framework for Fair Multilingual Cultural Evaluation of LLMs
by: Huang, Shulin, et al.
Published: (2025) -
CycleResearcher: Improving Automated Research via Automated Review
by: Weng, Yixuan, et al.
Published: (2024) -
AI Scientists Fail Without Strong Implementation Capability
by: Zhu, Minjun, et al.
Published: (2025)