GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
Fuente:
arXiv
Saved in:
| Main Authors: | Xiang, Yuxiao, Chen, Junchi, Jin, Zhenchao, Miao, Changtao, Yuan, Haojie, Chu, Qi, Gong, Tao, Yu, Nenghai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
by: Zhao, Jiawei, et al.
Published: (2025)
by: Zhao, Jiawei, et al.
Published: (2025)
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
Training-Free In-Context Forensic Chain for Image Manipulation Detection and Localization
by: Chen, Rui, et al.
Published: (2025)
by: Chen, Rui, et al.
Published: (2025)
LLMGuard: Guarding Against Unsafe LLM Behavior
by: Goyal, Shubh, et al.
Published: (2024)
by: Goyal, Shubh, et al.
Published: (2024)
PoseGuard: Pose-Guided Generation with Safety Guardrails
by: Wang, Kongxin, et al.
Published: (2025)
by: Wang, Kongxin, et al.
Published: (2025)
TraceGuard: Process-Guided Firewall against Reasoning Backdoors in Large Language Models
by: Guo, Zhen, et al.
Published: (2026)
by: Guo, Zhen, et al.
Published: (2026)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
by: Yuan, Lingzhi, et al.
Published: (2025)
by: Yuan, Lingzhi, et al.
Published: (2025)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
by: Shen, Guobin, et al.
Published: (2025)
by: Shen, Guobin, et al.
Published: (2025)
WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents
by: Chen, Yulin, et al.
Published: (2026)
by: Chen, Yulin, et al.
Published: (2026)
Friend or Foe Inside? Exploring In-Process Isolation to Maintain Memory Safety for Unsafe Rust
by: Gülmez, Merve, et al.
Published: (2023)
by: Gülmez, Merve, et al.
Published: (2023)
GLiNER Guard: Unified Encoder Family for Production LLM Safety and Privacy
by: Minko, Bogdan, et al.
Published: (2026)
by: Minko, Bogdan, et al.
Published: (2026)
Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset
by: Kang, Mintong, et al.
Published: (2025)
by: Kang, Mintong, et al.
Published: (2025)
Targeted Fuzzing for Unsafe Rust Code: Leveraging Selective Instrumentation
by: Paaßen, David, et al.
Published: (2025)
by: Paaßen, David, et al.
Published: (2025)
Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation
by: Wang, Haoyu, et al.
Published: (2024)
by: Wang, Haoyu, et al.
Published: (2024)
PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification
by: Gong, Guangyu, et al.
Published: (2026)
by: Gong, Guangyu, et al.
Published: (2026)
GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio
by: Zhu, Zhenhao, et al.
Published: (2026)
by: Zhu, Zhenhao, et al.
Published: (2026)
How to Steal Reasoning Without Reasoning Traces
by: Zhang, Tingwei, et al.
Published: (2026)
by: Zhang, Tingwei, et al.
Published: (2026)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
GIFDL: Generated Image Fluctuation Distortion Learning for Enhancing Steganographic Security
by: Wang, Xiangkun, et al.
Published: (2025)
by: Wang, Xiangkun, et al.
Published: (2025)
Silent Guardian: Protecting Text from Malicious Exploitation by Large Language Models
by: Zhao, Jiawei, et al.
Published: (2023)
by: Zhao, Jiawei, et al.
Published: (2023)
From Stealthy Data Fabrication to Unsafe Driving: Realistic Scenario Attacks on Collaborative Perception
by: Zhang, Qingzhao, et al.
Published: (2026)
by: Zhang, Qingzhao, et al.
Published: (2026)
Fast Summary-based Whole-program Analysis to Identify Unsafe Memory Accesses in Rust
by: Zhou, Jie, et al.
Published: (2023)
by: Zhou, Jie, et al.
Published: (2023)
On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts
by: Wu, Yixin, et al.
Published: (2023)
by: Wu, Yixin, et al.
Published: (2023)
A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations
by: Ye, Mang, et al.
Published: (2025)
by: Ye, Mang, et al.
Published: (2025)
AgentGuard: Repurposing Agentic Orchestrator for Safety Evaluation of Tool Orchestration
by: Chen, Jizhou, et al.
Published: (2025)
by: Chen, Jizhou, et al.
Published: (2025)
SQL Injection Jailbreak: A Structural Disaster of Large Language Models
by: Zhao, Jiawei, et al.
Published: (2024)
by: Zhao, Jiawei, et al.
Published: (2024)
Performance-lossless Black-box Model Watermarking
by: Zhao, Na, et al.
Published: (2023)
by: Zhao, Na, et al.
Published: (2023)
ProtoGuard-SL: Prototype Consistency Based Backdoor Defense for Vertical Split Learning
by: Shui, Yuhan, et al.
Published: (2026)
by: Shui, Yuhan, et al.
Published: (2026)
TraceGuard: Structured Multi-Dimensional Monitoring as a Collusion-Resistant Control Protocol
by: Nguyen, Khanh Linh, et al.
Published: (2026)
by: Nguyen, Khanh Linh, et al.
Published: (2026)
Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety
by: Zhang, Yuyou, et al.
Published: (2025)
by: Zhang, Yuyou, et al.
Published: (2025)
SandCell: Sandboxing Rust Beyond Unsafe Code
by: Zhang, Jialun, et al.
Published: (2025)
by: Zhang, Jialun, et al.
Published: (2025)
FASR: Automated Identification of Unsafe Control Actions in STPA
by: Dardik, Ian, et al.
Published: (2026)
by: Dardik, Ian, et al.
Published: (2026)
A Construction of Evolving $k$-threshold Secret Sharing Scheme over A Polynomial Ring
by: Cheng, Qi, et al.
Published: (2024)
by: Cheng, Qi, et al.
Published: (2024)
STEAD: Robust Provably Secure Linguistic Steganography with Diffusion Language Model
by: Qi, Yuang, et al.
Published: (2026)
by: Qi, Yuang, et al.
Published: (2026)
QGuard:Question-based Zero-shot Guard for Multi-modal LLM Safety
by: Lee, Taegyeong, et al.
Published: (2025)
by: Lee, Taegyeong, et al.
Published: (2025)
Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
by: Luo, Zeren, et al.
Published: (2025)
by: Luo, Zeren, et al.
Published: (2025)
MindGuard: Intrinsic Decision Inspection for Securing LLM Agents Against Metadata Poisoning
by: Wang, Zhiqiang, et al.
Published: (2025)
by: Wang, Zhiqiang, et al.
Published: (2025)
xIDS-EnsembleGuard: An Explainable Ensemble Learning-based Intrusion Detection System
by: Adil, Muhammad, et al.
Published: (2025)
by: Adil, Muhammad, et al.
Published: (2025)
GuardReasoner: Towards Reasoning-based LLM Safeguards
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
MSWasm: Soundly Enforcing Memory-Safe Execution of Unsafe Code
by: Michael, Alexandra E., et al.
Published: (2022)
by: Michael, Alexandra E., et al.
Published: (2022)
Similar Items
-
PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
by: Zhao, Jiawei, et al.
Published: (2025) -
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
by: Liu, Yue, et al.
Published: (2025) -
Training-Free In-Context Forensic Chain for Image Manipulation Detection and Localization
by: Chen, Rui, et al.
Published: (2025) -
LLMGuard: Guarding Against Unsafe LLM Behavior
by: Goyal, Shubh, et al.
Published: (2024) -
PoseGuard: Pose-Guided Generation with Safety Guardrails
by: Wang, Kongxin, et al.
Published: (2025)