LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ren, Qibing, Li, Hao, Liu, Dongrui, Xie, Zhanxu, Lu, Xiaoya, Qiao, Yu, Sha, Lei, Yan, Junchi, Ma, Lizhuang, Shao, Jing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion
von: Ren, Qibing, et al.
Veröffentlicht: (2024)
von: Ren, Qibing, et al.
Veröffentlicht: (2024)
When AI Agents Collude Online: Financial Fraud Risks by Collaborative LLM Agents on Social Platforms
von: Ren, Qibing, et al.
Veröffentlicht: (2025)
von: Ren, Qibing, et al.
Veröffentlicht: (2025)
When Autonomy Goes Rogue: Preparing for Risks of Multi-Agent Collusion in Social Systems
von: Ren, Qibing, et al.
Veröffentlicht: (2025)
von: Ren, Qibing, et al.
Veröffentlicht: (2025)
INFA-Guard: Mitigating Malicious Propagation via Infection-Aware Safeguarding in LLM-Based Multi-Agent Systems
von: Zhou, Yijin, et al.
Veröffentlicht: (2026)
von: Zhou, Yijin, et al.
Veröffentlicht: (2026)
X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Multi-Turn Jailbreaks without Compromising Usability
von: Lu, Xiaoya, et al.
Veröffentlicht: (2025)
von: Lu, Xiaoya, et al.
Veröffentlicht: (2025)
LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
von: Hu, Xuhao, et al.
Veröffentlicht: (2025)
von: Hu, Xuhao, et al.
Veröffentlicht: (2025)
IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
von: Lu, Xiaoya, et al.
Veröffentlicht: (2025)
von: Lu, Xiaoya, et al.
Veröffentlicht: (2025)
The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
VLSBench: Unveiling Visual Leakage in Multimodal Safety
von: Hu, Xuhao, et al.
Veröffentlicht: (2024)
von: Hu, Xuhao, et al.
Veröffentlicht: (2024)
Loop as a Bridge: Can Looped Transformers Truly Link Representation Space and Natural Language Outputs?
von: Chen, Guanxu, et al.
Veröffentlicht: (2026)
von: Chen, Guanxu, et al.
Veröffentlicht: (2026)
Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment
von: Li, Hao, et al.
Veröffentlicht: (2025)
von: Li, Hao, et al.
Veröffentlicht: (2025)
The Better Angels of Machine Personality: How Personality Relates to LLM Safety
von: Zhang, Jie, et al.
Veröffentlicht: (2024)
von: Zhang, Jie, et al.
Veröffentlicht: (2024)
ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
Valence-Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control
von: Sun, Lihao, et al.
Veröffentlicht: (2026)
von: Sun, Lihao, et al.
Veröffentlicht: (2026)
Handling Distribution Shifts on Graphs: An Invariance Perspective
von: Wu, Qitian, et al.
Veröffentlicht: (2022)
von: Wu, Qitian, et al.
Veröffentlicht: (2022)
LED-Merging: Mitigating Safety-Utility Conflicts in Model Merging with Location-Election-Disjoint
von: Ma, Qianli, et al.
Veröffentlicht: (2025)
von: Ma, Qianli, et al.
Veröffentlicht: (2025)
HomeGuard: VLM-based Embodied Safeguard for Identifying Contextual Risk in Household Task
von: Lu, Xiaoya, et al.
Veröffentlicht: (2026)
von: Lu, Xiaoya, et al.
Veröffentlicht: (2026)
Distribution Shift Alignment Helps LLMs Simulate Survey Response Distributions
von: Huang, Ji, et al.
Veröffentlicht: (2025)
von: Huang, Ji, et al.
Veröffentlicht: (2025)
Unified Batch Normalization: Identifying and Alleviating the Feature Condensation in Batch Normalization and a Unified Framework
von: Wang, Shaobo, et al.
Veröffentlicht: (2023)
von: Wang, Shaobo, et al.
Veröffentlicht: (2023)
Uncovering Gaps in How Humans and LLMs Interpret Subjective Language
von: Jones, Erik, et al.
Veröffentlicht: (2025)
von: Jones, Erik, et al.
Veröffentlicht: (2025)
RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
von: Yang, Jingyi, et al.
Veröffentlicht: (2025)
von: Yang, Jingyi, et al.
Veröffentlicht: (2025)
Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies
von: Liu, Chenruo, et al.
Veröffentlicht: (2025)
von: Liu, Chenruo, et al.
Veröffentlicht: (2025)
ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
Bounding Distributional Shifts in World Modeling through Novelty Detection
von: Jing, Eric, et al.
Veröffentlicht: (2025)
von: Jing, Eric, et al.
Veröffentlicht: (2025)
VISAT: Benchmarking Adversarial and Distribution Shift Robustness in Traffic Sign Recognition with Visual Attributes
von: Yu, Simon, et al.
Veröffentlicht: (2025)
von: Yu, Simon, et al.
Veröffentlicht: (2025)
TodoEvolve: Learning to Architect Agent Planning Systems
von: Liu, Jiaxi, et al.
Veröffentlicht: (2026)
von: Liu, Jiaxi, et al.
Veröffentlicht: (2026)
Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex
von: Yang, Zhonghao, et al.
Veröffentlicht: (2026)
von: Yang, Zhonghao, et al.
Veröffentlicht: (2026)
Do LLMs "know" internally when they follow instructions?
von: Heo, Juyeon, et al.
Veröffentlicht: (2024)
von: Heo, Juyeon, et al.
Veröffentlicht: (2024)
Uncovering the inherited vulnerability of electric distribution networks
von: Hartmann, Bálint, et al.
Veröffentlicht: (2024)
von: Hartmann, Bálint, et al.
Veröffentlicht: (2024)
CLA ‐ UNet : Convolution and Focused Linear Attention Fusion for Tumor Cell Nucleus Segmentation
von: Wei Guo, et al.
Veröffentlicht: (2025)
von: Wei Guo, et al.
Veröffentlicht: (2025)
VLLaVO: Mitigating Visual Gap through LLMs
von: Chen, Shuhao, et al.
Veröffentlicht: (2024)
von: Chen, Shuhao, et al.
Veröffentlicht: (2024)
3D Magnetic Field Reconstruction and Mapping with Physics-Informed Neural Networks
von: Yu, Haohan, et al.
Veröffentlicht: (2026)
von: Yu, Haohan, et al.
Veröffentlicht: (2026)
PRISM: Preference-Aware Influence Function Based Data Selection Method for Efficient Fine-Tuning
von: Lin, Qihao, et al.
Veröffentlicht: (2026)
von: Lin, Qihao, et al.
Veröffentlicht: (2026)
Adaptive Individual Uncertainty under Out-Of-Distribution Shift with Expert-Routed Conformal Prediction
von: Badkul, Amitesh, et al.
Veröffentlicht: (2025)
von: Badkul, Amitesh, et al.
Veröffentlicht: (2025)
Learning Calibrated Uncertainties for Domain Shift: A Distributionally Robust Learning Approach
von: Wang, Haoxuan, et al.
Veröffentlicht: (2020)
von: Wang, Haoxuan, et al.
Veröffentlicht: (2020)
Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
One RL to See Them All: Visual Triple Unified Reinforcement Learning
von: Ma, Yan, et al.
Veröffentlicht: (2025)
von: Ma, Yan, et al.
Veröffentlicht: (2025)
Why Perturbing Symbolic Music is Necessary: Fitting the Distribution of Never-used Notes through a Joint Probabilistic Diffusion Model
von: Liu, Shipei, et al.
Veröffentlicht: (2024)
von: Liu, Shipei, et al.
Veröffentlicht: (2024)
Uncovering Knowledge Gaps in Radiology Report Generation Models through Knowledge Graphs
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion
von: Ren, Qibing, et al.
Veröffentlicht: (2024) -
When AI Agents Collude Online: Financial Fraud Risks by Collaborative LLM Agents on Social Platforms
von: Ren, Qibing, et al.
Veröffentlicht: (2025) -
When Autonomy Goes Rogue: Preparing for Risks of Multi-Agent Collusion in Social Systems
von: Ren, Qibing, et al.
Veröffentlicht: (2025) -
INFA-Guard: Mitigating Malicious Propagation via Infection-Aware Safeguarding in LLM-Based Multi-Agent Systems
von: Zhou, Yijin, et al.
Veröffentlicht: (2026) -
X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Multi-Turn Jailbreaks without Compromising Usability
von: Lu, Xiaoya, et al.
Veröffentlicht: (2025)