Cat-DPO: Category-Adaptive Safety Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Tiankai, Nian, Yi, Li, Xinyuan, Xu, Ruiyao, Ding, Kaize, Zhao, Yue |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
No Attacker Needed: Unintentional Cross-User Contamination in Shared-State LLM Agents
von: Yang, Tiankai, et al.
Veröffentlicht: (2026)
von: Yang, Tiankai, et al.
Veröffentlicht: (2026)
GNN-as-Judge: Unleashing the Power of LLMs for Graph Learning with GNN Feedback
von: Xu, Ruiyao, et al.
Veröffentlicht: (2026)
von: Xu, Ruiyao, et al.
Veröffentlicht: (2026)
AD-LLM: Benchmarking Large Language Models for Anomaly Detection
von: Yang, Tiankai, et al.
Veröffentlicht: (2024)
von: Yang, Tiankai, et al.
Veröffentlicht: (2024)
CoAct: Co-Active LLM Preference Learning with Human-AI Synergy
von: Xu, Ruiyao, et al.
Veröffentlicht: (2026)
von: Xu, Ruiyao, et al.
Veröffentlicht: (2026)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment
von: Hu, Mengxuan, et al.
Veröffentlicht: (2026)
von: Hu, Mengxuan, et al.
Veröffentlicht: (2026)
Synthetic Clinical Notes for Rare ICD Codes: A Data-Centric Framework for Long-Tail Medical Coding
von: Vo, Truong, et al.
Veröffentlicht: (2025)
von: Vo, Truong, et al.
Veröffentlicht: (2025)
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
von: Lee, Andrew, et al.
Veröffentlicht: (2024)
von: Lee, Andrew, et al.
Veröffentlicht: (2024)
Empowering Large Language Models for Textual Data Augmentation
von: Li, Yichuan, et al.
Veröffentlicht: (2024)
von: Li, Yichuan, et al.
Veröffentlicht: (2024)
MixDPO: Modeling Preference Strength for Pluralistic Alignment
von: Imai, Saki, et al.
Veröffentlicht: (2026)
von: Imai, Saki, et al.
Veröffentlicht: (2026)
Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
STAR-S: Improving Safety Alignment through Self-Taught Reasoning on Safety Rules
von: Wu, Di, et al.
Veröffentlicht: (2026)
von: Wu, Di, et al.
Veröffentlicht: (2026)
Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences
von: Pattnaik, Pulkit, et al.
Veröffentlicht: (2024)
von: Pattnaik, Pulkit, et al.
Veröffentlicht: (2024)
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
von: He, Chaoyue, et al.
Veröffentlicht: (2026)
von: He, Chaoyue, et al.
Veröffentlicht: (2026)
StealthRank: LLM Ranking Manipulation via Stealthy Prompt Optimization
von: Tang, Yiming, et al.
Veröffentlicht: (2025)
von: Tang, Yiming, et al.
Veröffentlicht: (2025)
An Empirical Study of SFT-DPO Interaction and Parameterization in Small Language Models
von: Feng, Yuming, et al.
Veröffentlicht: (2026)
von: Feng, Yuming, et al.
Veröffentlicht: (2026)
The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval
von: Tong, Zekai, et al.
Veröffentlicht: (2026)
von: Tong, Zekai, et al.
Veröffentlicht: (2026)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
von: Zhang, Zhengze, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengze, et al.
Veröffentlicht: (2025)
Self-Guided Defense: Adaptive Safety Alignment for Reasoning Models via Synthesized Guidelines
von: Wang, Yuhang, et al.
Veröffentlicht: (2025)
von: Wang, Yuhang, et al.
Veröffentlicht: (2025)
Don't Let It Hallucinate: Premise Verification via Retrieval-Augmented Logical Reasoning
von: Qin, Yuehan, et al.
Veröffentlicht: (2025)
von: Qin, Yuehan, et al.
Veröffentlicht: (2025)
APLe: Token-Wise Adaptive for Multi-Modal Prompt Learning
von: Cao, Guiming, et al.
Veröffentlicht: (2024)
von: Cao, Guiming, et al.
Veröffentlicht: (2024)
Safety Alignment via Constrained Knowledge Unlearning
von: Shi, Zesheng, et al.
Veröffentlicht: (2025)
von: Shi, Zesheng, et al.
Veröffentlicht: (2025)
SmurfCat at PAN 2024 TextDetox: Alignment of Multilingual Transformers for Text Detoxification
von: Rykov, Elisei, et al.
Veröffentlicht: (2024)
von: Rykov, Elisei, et al.
Veröffentlicht: (2024)
Aligning Large Language Models with Counterfactual DPO
von: Butcher, Bradley
Veröffentlicht: (2024)
von: Butcher, Bradley
Veröffentlicht: (2024)
Safety Is Not Universal: The Selective Safety Trap in LLM Alignment
von: Brito, Iago Alves, et al.
Veröffentlicht: (2026)
von: Brito, Iago Alves, et al.
Veröffentlicht: (2026)
MedCalc-Eval and MedCalc-Env: Advancing Medical Calculation Capabilities of Large Language Models
von: Mao, Kangkun, et al.
Veröffentlicht: (2025)
von: Mao, Kangkun, et al.
Veröffentlicht: (2025)
AD-AGENT: A Multi-agent Framework for End-to-end Anomaly Detection
von: Yang, Tiankai, et al.
Veröffentlicht: (2025)
von: Yang, Tiankai, et al.
Veröffentlicht: (2025)
Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M
von: Pant, Piyush
Veröffentlicht: (2025)
von: Pant, Piyush
Veröffentlicht: (2025)
Towards Inference-time Category-wise Safety Steering for Large Language Models
von: Bhattacharjee, Amrita, et al.
Veröffentlicht: (2024)
von: Bhattacharjee, Amrita, et al.
Veröffentlicht: (2024)
Rethinking DPO: The Role of Rejected Responses in Preference Misalignment
von: Cho, Jay Hyeon, et al.
Veröffentlicht: (2025)
von: Cho, Jay Hyeon, et al.
Veröffentlicht: (2025)
DPO Meets PPO: Reinforced Token Optimization for RLHF
von: Zhong, Han, et al.
Veröffentlicht: (2024)
von: Zhong, Han, et al.
Veröffentlicht: (2024)
When Only the Final Text Survives: Implicit Execution Tracing for Multi-Agent Attribution
von: Nian, Yi, et al.
Veröffentlicht: (2026)
von: Nian, Yi, et al.
Veröffentlicht: (2026)
Reasoning Pattern Alignment Merging for Adaptive Reasoning
von: Zhong, Zhaofeng, et al.
Veröffentlicht: (2026)
von: Zhong, Zhaofeng, et al.
Veröffentlicht: (2026)
SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
von: Li, Hao, et al.
Veröffentlicht: (2026)
von: Li, Hao, et al.
Veröffentlicht: (2026)
RobustFT: Robust Supervised Fine-tuning for Large Language Models under Noisy Response
von: Luo, Junyu, et al.
Veröffentlicht: (2024)
von: Luo, Junyu, et al.
Veröffentlicht: (2024)
Towards Context-Invariant Safety Alignment for Large Language Models
von: Wang, Yixu, et al.
Veröffentlicht: (2026)
von: Wang, Yixu, et al.
Veröffentlicht: (2026)
Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
von: Zhang, Jingyu, et al.
Veröffentlicht: (2024)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2024)
Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective
von: Feng, Duanyu, et al.
Veröffentlicht: (2024)
von: Feng, Duanyu, et al.
Veröffentlicht: (2024)
sDPO: Don't Use Your Data All at Once
von: Kim, Dahyun, et al.
Veröffentlicht: (2024)
von: Kim, Dahyun, et al.
Veröffentlicht: (2024)
RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models
von: Khaki, Saeed, et al.
Veröffentlicht: (2024)
von: Khaki, Saeed, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
No Attacker Needed: Unintentional Cross-User Contamination in Shared-State LLM Agents
von: Yang, Tiankai, et al.
Veröffentlicht: (2026) -
GNN-as-Judge: Unleashing the Power of LLMs for Graph Learning with GNN Feedback
von: Xu, Ruiyao, et al.
Veröffentlicht: (2026) -
AD-LLM: Benchmarking Large Language Models for Anomaly Detection
von: Yang, Tiankai, et al.
Veröffentlicht: (2024) -
CoAct: Co-Active LLM Preference Learning with Human-AI Synergy
von: Xu, Ruiyao, et al.
Veröffentlicht: (2026) -
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
von: Wu, Junkang, et al.
Veröffentlicht: (2024)