ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Zhexin, Lu, Yida, Ma, Jingyuan, Zhang, Di, Li, Rui, Ke, Pei, Sun, Hao, Sha, Lei, Sui, Zhifang, Wang, Hongning, Huang, Minlie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Agent-SafetyBench: Evaluating the Safety of LLM Agents
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
von: Cui, Shiyao, et al.
Veröffentlicht: (2025)
von: Cui, Shiyao, et al.
Veröffentlicht: (2025)
From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
von: Zhang, Zhexin, et al.
Veröffentlicht: (2023)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2023)
Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints
von: Yang, Junxiao, et al.
Veröffentlicht: (2025)
von: Yang, Junxiao, et al.
Veröffentlicht: (2025)
Be a Multitude to Itself: A Prompt Evolution Framework for Red Teaming
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
LongSafety: Evaluating Long-Context Safety of Large Language Models
von: Lu, Yida, et al.
Veröffentlicht: (2025)
von: Lu, Yida, et al.
Veröffentlicht: (2025)
AISafetyLab: A Comprehensive Framework for AI Safety Evaluation and Improvement
von: Zhang, Zhexin, et al.
Veröffentlicht: (2025)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2025)
Plug-and-Play Training Framework for Preference Optimization
von: Ma, Jingyuan, et al.
Veröffentlicht: (2024)
von: Ma, Jingyuan, et al.
Veröffentlicht: (2024)
Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!
von: Zhang, Zhexin, et al.
Veröffentlicht: (2025)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2025)
The Missing Half: Unveiling Training-time Implicit Safety Risks Beyond Deployment
von: Zhang, Zhexin, et al.
Veröffentlicht: (2026)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2026)
Language Model Decoding as Direct Metrics Optimization
von: Ji, Haozhe, et al.
Veröffentlicht: (2023)
von: Ji, Haozhe, et al.
Veröffentlicht: (2023)
ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity
von: Cui, Shiyao, et al.
Veröffentlicht: (2025)
von: Cui, Shiyao, et al.
Veröffentlicht: (2025)
SafetyBench: Evaluating the Safety of Large Language Models
von: Zhang, Zhexin, et al.
Veröffentlicht: (2023)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2023)
HauntAttack: When Attack Follows Reasoning as a Shadow
von: Ma, Jingyuan, et al.
Veröffentlicht: (2025)
von: Ma, Jingyuan, et al.
Veröffentlicht: (2025)
Black-Box Prompt Optimization: Aligning Large Language Models without Model Training
von: Cheng, Jiale, et al.
Veröffentlicht: (2023)
von: Cheng, Jiale, et al.
Veröffentlicht: (2023)
Seeker: Towards Exception Safety Code Generation with Intermediate Language Agents Framework
von: Zhang, Xuanming, et al.
Veröffentlicht: (2024)
von: Zhang, Xuanming, et al.
Veröffentlicht: (2024)
AutoDetect: Towards a Unified Framework for Automated Weakness Detection in Large Language Models
von: Cheng, Jiale, et al.
Veröffentlicht: (2024)
von: Cheng, Jiale, et al.
Veröffentlicht: (2024)
Learning Task Decomposition to Assist Humans in Competitive Programming
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
How Far are LLMs from Being Our Digital Twins? A Benchmark for Persona-Based Behavior Chain Simulation
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
Grounding LLMs in Scientific Discovery via Embodied Actions
von: Zhang, Bo, et al.
Veröffentlicht: (2026)
von: Zhang, Bo, et al.
Veröffentlicht: (2026)
LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
von: Yang, Junxiao, et al.
Veröffentlicht: (2026)
von: Yang, Junxiao, et al.
Veröffentlicht: (2026)
SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
von: Li, Zheng, et al.
Veröffentlicht: (2025)
von: Li, Zheng, et al.
Veröffentlicht: (2025)
How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study
von: Zhang, Zhexin, et al.
Veröffentlicht: (2025)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2025)
Large Language Models Struggle with Unreasonability in Math Problems
von: Ma, Jingyuan, et al.
Veröffentlicht: (2024)
von: Ma, Jingyuan, et al.
Veröffentlicht: (2024)
VPO: Aligning Text-to-Video Generation Models with Prompt Optimization
von: Cheng, Jiale, et al.
Veröffentlicht: (2025)
von: Cheng, Jiale, et al.
Veröffentlicht: (2025)
IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation
von: Wen, Bosi, et al.
Veröffentlicht: (2026)
von: Wen, Bosi, et al.
Veröffentlicht: (2026)
StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation
von: Wen, Bosi, et al.
Veröffentlicht: (2025)
von: Wen, Bosi, et al.
Veröffentlicht: (2025)
Towards Harmonized Uncertainty Estimation for Large Language Models
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
AlignBench: Benchmarking Chinese Alignment of Large Language Models
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
Towards Efficient Exact Optimization of Language Model Alignment
von: Ji, Haozhe, et al.
Veröffentlicht: (2024)
von: Ji, Haozhe, et al.
Veröffentlicht: (2024)
JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual Steering
von: Chen, Renmiao, et al.
Veröffentlicht: (2025)
von: Chen, Renmiao, et al.
Veröffentlicht: (2025)
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
von: Cheng, Jiale, et al.
Veröffentlicht: (2024)
von: Cheng, Jiale, et al.
Veröffentlicht: (2024)
A High‐Energy and High‐Safety Bisolvent‐in‐Salt Semisolid Flow Battery
von: Xinyi Zou, et al.
Veröffentlicht: (2026)
von: Xinyi Zou, et al.
Veröffentlicht: (2026)
FraudShield: Knowledge Graph Empowered Defense for LLMs against Fraud Attacks
von: Xu, Naen, et al.
Veröffentlicht: (2026)
von: Xu, Naen, et al.
Veröffentlicht: (2026)
TeachBench: A Syllabus-Grounded Framework for Evaluating Teaching Ability in Large Language Models
von: Li, Zheng, et al.
Veröffentlicht: (2026)
von: Li, Zheng, et al.
Veröffentlicht: (2026)
Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
von: Ma, Wenhan, et al.
Veröffentlicht: (2025)
von: Ma, Wenhan, et al.
Veröffentlicht: (2025)
ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Agent-SafetyBench: Evaluating the Safety of LLM Agents
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024) -
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
von: Cui, Shiyao, et al.
Veröffentlicht: (2025) -
From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024) -
Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
von: Zhang, Zhexin, et al.
Veröffentlicht: (2023) -
Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints
von: Yang, Junxiao, et al.
Veröffentlicht: (2025)