The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Rui, Li, Hongwei, Shen, Yun, Shen, Xinyue, Jiang, Wenbo, Xu, Guowen, Liu, Yang, Backes, Michael, Zhang, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Instruction Backdoor Attacks Against Customized LLMs
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTs
von: Shen, Xinyue, et al.
Veröffentlicht: (2025)
von: Shen, Xinyue, et al.
Veröffentlicht: (2025)
BadTemplate: A Training-Free Backdoor Attack via Chat Template Against Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2026)
von: Wang, Zihan, et al.
Veröffentlicht: (2026)
The Ripple Effect: On Unforeseen Complications of Backdoor Attacks
von: Zhang, Rui, et al.
Veröffentlicht: (2025)
von: Zhang, Rui, et al.
Veröffentlicht: (2025)
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
von: Shen, Xinyue, et al.
Veröffentlicht: (2023)
von: Shen, Xinyue, et al.
Veröffentlicht: (2023)
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
von: Jiang, Yukun, et al.
Veröffentlicht: (2026)
von: Jiang, Yukun, et al.
Veröffentlicht: (2026)
Pop Quiz Attack: Black-box Membership Inference Attacks Against Large Language Models
von: Chen, Zeyuan, et al.
Veröffentlicht: (2026)
von: Chen, Zeyuan, et al.
Veröffentlicht: (2026)
Voice Jailbreak Attacks Against GPT-4o
von: Shen, Xinyue, et al.
Veröffentlicht: (2024)
von: Shen, Xinyue, et al.
Veröffentlicht: (2024)
Prompt Stealing Attacks Against Text-to-Image Generation Models
von: Shen, Xinyue, et al.
Veröffentlicht: (2023)
von: Shen, Xinyue, et al.
Veröffentlicht: (2023)
Robustness Over Time: Understanding Adversarial Examples' Effectiveness on Longitudinal Versions of Large Language Models
von: Liu, Yugeng, et al.
Veröffentlicht: (2023)
von: Liu, Yugeng, et al.
Veröffentlicht: (2023)
ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
von: Chu, Junjie, et al.
Veröffentlicht: (2026)
von: Chu, Junjie, et al.
Veröffentlicht: (2026)
On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts
von: Wu, Yixin, et al.
Veröffentlicht: (2023)
von: Wu, Yixin, et al.
Veröffentlicht: (2023)
MGTBench: Benchmarking Machine-Generated Text Detection
von: He, Xinlei, et al.
Veröffentlicht: (2023)
von: He, Xinlei, et al.
Veröffentlicht: (2023)
Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model Evolution
von: Wu, Yixin, et al.
Veröffentlicht: (2024)
von: Wu, Yixin, et al.
Veröffentlicht: (2024)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
Combinational Backdoor Attack against Customized Text-to-Image Models
von: Jiang, Wenbo, et al.
Veröffentlicht: (2024)
von: Jiang, Wenbo, et al.
Veröffentlicht: (2024)
HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns
von: Shen, Xinyue, et al.
Veröffentlicht: (2025)
von: Shen, Xinyue, et al.
Veröffentlicht: (2025)
$\texttt{ModSCAN}$: Measuring Stereotypical Bias in Large Vision-Language Models from Vision and Language Modalities
von: Jiang, Yukun, et al.
Veröffentlicht: (2024)
von: Jiang, Yukun, et al.
Veröffentlicht: (2024)
"Humans welcome to observe": A First Look at the Agent Social Network Moltbook
von: Jiang, Yukun, et al.
Veröffentlicht: (2026)
von: Jiang, Yukun, et al.
Veröffentlicht: (2026)
BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning
von: Yang, Ziqing, et al.
Veröffentlicht: (2026)
von: Yang, Ziqing, et al.
Veröffentlicht: (2026)
The Challenge of Identifying the Origin of Black-Box Large Language Models
von: Yang, Ziqing, et al.
Veröffentlicht: (2025)
von: Yang, Ziqing, et al.
Veröffentlicht: (2025)
Hidden Tail: Adversarial Image Causing Stealthy Resource Consumption in Vision-Language Models
von: Zhang, Rui, et al.
Veröffentlicht: (2025)
von: Zhang, Rui, et al.
Veröffentlicht: (2025)
Synthetic Artifact Auditing: Tracing LLM-Generated Synthetic Data Usage in Downstream Applications
von: Wu, Yixin, et al.
Veröffentlicht: (2025)
von: Wu, Yixin, et al.
Veröffentlicht: (2025)
BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
von: Zhang, Yage, et al.
Veröffentlicht: (2026)
von: Zhang, Yage, et al.
Veröffentlicht: (2026)
Efficient Data-Free Model Stealing with Label Diversity
von: Liu, Yiyong, et al.
Veröffentlicht: (2024)
von: Liu, Yiyong, et al.
Veröffentlicht: (2024)
Peering Behind the Shield: Guardrail Identification in Large Language Models
von: Yang, Ziqing, et al.
Veröffentlicht: (2025)
von: Yang, Ziqing, et al.
Veröffentlicht: (2025)
Do Fine-Tuned LLMs Understand Vulnerabilities? An Investigation into the Semantic Trap
von: Huang, Feiyang, et al.
Veröffentlicht: (2026)
von: Huang, Feiyang, et al.
Veröffentlicht: (2026)
MPMA: Preference Manipulation Attack Against Model Context Protocol
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
Understanding Data Importance in Machine Learning Attacks: Does Valuable Data Pose Greater Harm?
von: Wen, Rui, et al.
Veröffentlicht: (2024)
von: Wen, Rui, et al.
Veröffentlicht: (2024)
Differentially Private Subspace Fine-Tuning for Large Language Models
von: Zheng, Lele, et al.
Veröffentlicht: (2026)
von: Zheng, Lele, et al.
Veröffentlicht: (2026)
Composite Backdoor Attacks Against Large Language Models
von: Huang, Hai, et al.
Veröffentlicht: (2023)
von: Huang, Hai, et al.
Veröffentlicht: (2023)
Black-Box Skill Stealing Attack from Proprietary LLM Agents: An Empirical Study
von: Wang, Zihan, et al.
Veröffentlicht: (2026)
von: Wang, Zihan, et al.
Veröffentlicht: (2026)
Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
OnePath: Efficient and Privacy-Preserving Decision Tree Inference in the Cloud
von: Yuan, Shuai, et al.
Veröffentlicht: (2024)
von: Yuan, Shuai, et al.
Veröffentlicht: (2024)
JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring
von: Chu, Junjie, et al.
Veröffentlicht: (2025)
von: Chu, Junjie, et al.
Veröffentlicht: (2025)
Excessive Reasoning Attack on Reasoning LLMs
von: Si, Wai Man, et al.
Veröffentlicht: (2025)
von: Si, Wai Man, et al.
Veröffentlicht: (2025)
Understanding LLM Behavior When Encountering User-Supplied Harmful Content in Harmless Tasks
von: Chu, Junjie, et al.
Veröffentlicht: (2026)
von: Chu, Junjie, et al.
Veröffentlicht: (2026)
Backdoor Attacks against Image-to-Image Networks
von: Jiang, Wenbo, et al.
Veröffentlicht: (2024)
von: Jiang, Wenbo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Instruction Backdoor Attacks Against Customized LLMs
von: Zhang, Rui, et al.
Veröffentlicht: (2024) -
When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTs
von: Shen, Xinyue, et al.
Veröffentlicht: (2025) -
BadTemplate: A Training-Free Backdoor Attack via Chat Template Against Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2026) -
The Ripple Effect: On Unforeseen Complications of Backdoor Attacks
von: Zhang, Rui, et al.
Veröffentlicht: (2025) -
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
von: Shen, Xinyue, et al.
Veröffentlicht: (2023)