Picky LLMs and Unreliable RMs: An Empirical Study on Safety Alignment after Instruction Tuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Guanlin, Chen, Kangjie, Guo, Shangwei, Zhang, Jie, Qiu, Han, Zhang, Chao, Wang, Guoyin, Zhang, Tianwei, Li, Jiwei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Warfare:Breaking the Watermark Protection of AI-Generated Content
von: Li, Guanlin, et al.
Veröffentlicht: (2023)
von: Li, Guanlin, et al.
Veröffentlicht: (2023)
Instruction Tuning for Large Language Models: A Survey
von: Zhang, Shengyu, et al.
Veröffentlicht: (2023)
von: Zhang, Shengyu, et al.
Veröffentlicht: (2023)
Reinforcement Learning Enhanced LLMs: A Survey
von: Wang, Shuhe, et al.
Veröffentlicht: (2024)
von: Wang, Shuhe, et al.
Veröffentlicht: (2024)
Fingerprinting Image-to-Image Generative Adversarial Networks
von: Li, Guanlin, et al.
Veröffentlicht: (2021)
von: Li, Guanlin, et al.
Veröffentlicht: (2021)
Similarity-based Neighbor Selection for Graph LLMs
von: Li, Rui, et al.
Veröffentlicht: (2024)
von: Li, Rui, et al.
Veröffentlicht: (2024)
ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign Users
von: Li, Guanlin, et al.
Veröffentlicht: (2024)
von: Li, Guanlin, et al.
Veröffentlicht: (2024)
Template-assisted Contrastive Learning of Task-oriented Dialogue Sentence Embeddings
von: Oh, Minsik, et al.
Veröffentlicht: (2023)
von: Oh, Minsik, et al.
Veröffentlicht: (2023)
Inference-time Alignment via Sparse Junction Steering
von: Hu, Runyi, et al.
Veröffentlicht: (2026)
von: Hu, Runyi, et al.
Veröffentlicht: (2026)
Towards Building the Federated GPT: Federated Instruction Tuning
von: Zhang, Jianyi, et al.
Veröffentlicht: (2023)
von: Zhang, Jianyi, et al.
Veröffentlicht: (2023)
Course-Correction: Safety Alignment Using Synthetic Preferences
von: Xu, Rongwu, et al.
Veröffentlicht: (2024)
von: Xu, Rongwu, et al.
Veröffentlicht: (2024)
Cross-Task Defense: Instruction-Tuning LLMs for Content Safety
von: Fu, Yu, et al.
Veröffentlicht: (2024)
von: Fu, Yu, et al.
Veröffentlicht: (2024)
Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures
von: Su, Yanghao, et al.
Veröffentlicht: (2026)
von: Su, Yanghao, et al.
Veröffentlicht: (2026)
PK-ICR: Persona-Knowledge Interactive Context Retrieval for Grounded Dialogue
von: Oh, Minsik, et al.
Veröffentlicht: (2023)
von: Oh, Minsik, et al.
Veröffentlicht: (2023)
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
von: Ji, Jiaming, et al.
Veröffentlicht: (2024)
von: Ji, Jiaming, et al.
Veröffentlicht: (2024)
Understanding the Dark Side of LLMs' Intrinsic Self-Correction
von: Zhang, Qingjie, et al.
Veröffentlicht: (2024)
von: Zhang, Qingjie, et al.
Veröffentlicht: (2024)
The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs
von: Zhang, Xing, et al.
Veröffentlicht: (2026)
von: Zhang, Xing, et al.
Veröffentlicht: (2026)
Parameter Efficient Instruction Tuning: An Empirical Study
von: He, Pengfei
Veröffentlicht: (2024)
von: He, Pengfei
Veröffentlicht: (2024)
Speculating LLMs' Chinese Training Data Pollution from Their Tokens
von: Zhang, Qingjie, et al.
Veröffentlicht: (2025)
von: Zhang, Qingjie, et al.
Veröffentlicht: (2025)
DSB: Dynamic Sliding Block Scheduling for Diffusion LLMs
von: Luo, Lizhuo, et al.
Veröffentlicht: (2026)
von: Luo, Lizhuo, et al.
Veröffentlicht: (2026)
Robust-Wide: Robust Watermarking against Instruction-driven Image Editing
von: Hu, Runyi, et al.
Veröffentlicht: (2024)
von: Hu, Runyi, et al.
Veröffentlicht: (2024)
Sample Design Engineering: An Empirical Study of What Makes Good Downstream Fine-Tuning Samples for LLMs
von: Guo, Biyang, et al.
Veröffentlicht: (2024)
von: Guo, Biyang, et al.
Veröffentlicht: (2024)
Beyond Retrieval: Improving Evidence Quality for LLM-based Multimodal Fact-Checking
von: Ou, Haoran, et al.
Veröffentlicht: (2025)
von: Ou, Haoran, et al.
Veröffentlicht: (2025)
Does Instruction Tuning Make LLMs More Consistent?
von: Fierro, Constanza, et al.
Veröffentlicht: (2024)
von: Fierro, Constanza, et al.
Veröffentlicht: (2024)
PRIME: Protect Your Videos From Malicious Editing
von: Li, Guanlin, et al.
Veröffentlicht: (2024)
von: Li, Guanlin, et al.
Veröffentlicht: (2024)
BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning
von: Zhang, Junbo, et al.
Veröffentlicht: (2026)
von: Zhang, Junbo, et al.
Veröffentlicht: (2026)
Data Diversity Matters for Robust Instruction Tuning
von: Bukharin, Alexander, et al.
Veröffentlicht: (2023)
von: Bukharin, Alexander, et al.
Veröffentlicht: (2023)
Packing Analysis: Packing Is More Appropriate for Large Models or Datasets in Supervised Fine-tuning
von: Wang, Shuhe, et al.
Veröffentlicht: (2024)
von: Wang, Shuhe, et al.
Veröffentlicht: (2024)
FinSphere, a Real-Time Stock Analysis Agent Powered by Instruction-Tuned LLMs and Domain Tools
von: Han, Shijie, et al.
Veröffentlicht: (2025)
von: Han, Shijie, et al.
Veröffentlicht: (2025)
OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety
von: Liu, Chuang, et al.
Veröffentlicht: (2024)
von: Liu, Chuang, et al.
Veröffentlicht: (2024)
Facilitating Multi-turn Function Calling for LLMs via Compositional Instruction Tuning
von: Chen, Mingyang, et al.
Veröffentlicht: (2024)
von: Chen, Mingyang, et al.
Veröffentlicht: (2024)
SuperMark: Robust and Training-free Image Watermarking via Diffusion-based Super-Resolution
von: Hu, Runyi, et al.
Veröffentlicht: (2024)
von: Hu, Runyi, et al.
Veröffentlicht: (2024)
VideoShield: Regulating Diffusion-based Video Generation Models via Watermarking
von: Hu, Runyi, et al.
Veröffentlicht: (2025)
von: Hu, Runyi, et al.
Veröffentlicht: (2025)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
von: Wu, Xuansheng, et al.
Veröffentlicht: (2023)
von: Wu, Xuansheng, et al.
Veröffentlicht: (2023)
DeepSweep: An Evaluation Framework for Mitigating DNN Backdoor Attacks using Data Augmentation
von: Qiu, Han, et al.
Veröffentlicht: (2020)
von: Qiu, Han, et al.
Veröffentlicht: (2020)
Data Selection for Multi-turn Dialogue Instruction Tuning
von: Li, Bo, et al.
Veröffentlicht: (2026)
von: Li, Bo, et al.
Veröffentlicht: (2026)
Span-level Emotion-Cause-Category Triplet Extraction with Instruction Tuning LLMs and Data Augmentation
von: Li, Xiangju, et al.
Veröffentlicht: (2025)
von: Li, Xiangju, et al.
Veröffentlicht: (2025)
Tuning LLMs with Contrastive Alignment Instructions for Machine Translation in Unseen, Low-resource Languages
von: Mao, Zhuoyuan, et al.
Veröffentlicht: (2024)
von: Mao, Zhuoyuan, et al.
Veröffentlicht: (2024)
Tool Preferences in Agentic LLMs are Unreliable
von: Faghih, Kazem, et al.
Veröffentlicht: (2025)
von: Faghih, Kazem, et al.
Veröffentlicht: (2025)
Beware of Your Po! Measuring and Mitigating AI Safety Risks in Role-Play Fine-Tuning of LLMs
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Warfare:Breaking the Watermark Protection of AI-Generated Content
von: Li, Guanlin, et al.
Veröffentlicht: (2023) -
Instruction Tuning for Large Language Models: A Survey
von: Zhang, Shengyu, et al.
Veröffentlicht: (2023) -
Reinforcement Learning Enhanced LLMs: A Survey
von: Wang, Shuhe, et al.
Veröffentlicht: (2024) -
Fingerprinting Image-to-Image Generative Adversarial Networks
von: Li, Guanlin, et al.
Veröffentlicht: (2021) -
Similarity-based Neighbor Selection for Graph LLMs
von: Li, Rui, et al.
Veröffentlicht: (2024)