MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Gu, Tianle, Zhou, Zeyang, Huang, Kexin, Liang, Dandan, Wang, Yixu, Zhao, Haiquan, Yao, Yuanqi, Qiao, Xingge, Wang, Keqing, Yang, Yujiu, Teng, Yan, Qiao, Yu, Wang, Yingchun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Probing the Robustness of Large Language Models Safety to Latent Perturbations
por: Gu, Tianle, et al.
Publicado: (2025)
por: Gu, Tianle, et al.
Publicado: (2025)
HoneypotNet: Backdoor Attacks Against Model Extraction
por: Wang, Yixu, et al.
Publicado: (2025)
por: Wang, Yixu, et al.
Publicado: (2025)
Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking
por: Gu, Tianle, et al.
Publicado: (2025)
por: Gu, Tianle, et al.
Publicado: (2025)
OpenRT: An Open-Source Red Teaming Framework for Multimodal LLMs
por: Wang, Xin, et al.
Publicado: (2026)
por: Wang, Xin, et al.
Publicado: (2026)
StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic Data
por: Wang, Yixu, et al.
Publicado: (2025)
por: Wang, Yixu, et al.
Publicado: (2025)
MorphMark: Flexible Adaptive Watermarking for Large Language Models
por: Wang, Zongqi, et al.
Publicado: (2025)
por: Wang, Zongqi, et al.
Publicado: (2025)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
por: Chen, Yunhao, et al.
Publicado: (2025)
por: Chen, Yunhao, et al.
Publicado: (2025)
A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos
por: Yao, Yang, et al.
Publicado: (2025)
por: Yao, Yang, et al.
Publicado: (2025)
Contrastive Learning for Continuous Touch-Based Authentication
por: Qiao, Mengyu, et al.
Publicado: (2025)
por: Qiao, Mengyu, et al.
Publicado: (2025)
DETOUR: A Practical Backdoor Attack against Object Detection
por: Liu, Dazhuang, et al.
Publicado: (2026)
por: Liu, Dazhuang, et al.
Publicado: (2026)
Agent Safety Alignment via Reinforcement Learning
por: Sha, Zeyang, et al.
Publicado: (2025)
por: Sha, Zeyang, et al.
Publicado: (2025)
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
por: Ying, Zonghao, et al.
Publicado: (2024)
por: Ying, Zonghao, et al.
Publicado: (2024)
Safety Geometry Collapse in Multimodal LLMs and Adaptive Drift Correction
por: Guo, Jiahe, et al.
Publicado: (2026)
por: Guo, Jiahe, et al.
Publicado: (2026)
USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models
por: Zheng, Baolin, et al.
Publicado: (2025)
por: Zheng, Baolin, et al.
Publicado: (2025)
ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety
por: Wang, Kun, et al.
Publicado: (2026)
por: Wang, Kun, et al.
Publicado: (2026)
UniMark: Artificial Intelligence Generated Content Identification Toolkit
por: Li, Meilin, et al.
Publicado: (2025)
por: Li, Meilin, et al.
Publicado: (2025)
Graph Learning Across Data Silos
por: Zhang, Xiang, et al.
Publicado: (2023)
por: Zhang, Xiang, et al.
Publicado: (2023)
ProvX: Generating Counterfactual-Driven Attack Explanations for Provenance-Based Detection
por: Wu, Weiheng, et al.
Publicado: (2025)
por: Wu, Weiheng, et al.
Publicado: (2025)
When Safety Becomes a Vulnerability: Exploiting LLM Alignment Homogeneity for Transferable Blocking in RAG
por: Li, Junchen, et al.
Publicado: (2026)
por: Li, Junchen, et al.
Publicado: (2026)
SmartInv: Multimodal Learning for Smart Contract Invariant Inference
por: Wang, Sally Junsong, et al.
Publicado: (2024)
por: Wang, Sally Junsong, et al.
Publicado: (2024)
The Shadow of Fraud: The Emerging Danger of AI-powered Social Engineering and its Possible Cure
por: Yu, Jingru, et al.
Publicado: (2024)
por: Yu, Jingru, et al.
Publicado: (2024)
Evaluating Selective Encryption Against Gradient Inversion Attacks
por: Gu, Jiajun, et al.
Publicado: (2025)
por: Gu, Jiajun, et al.
Publicado: (2025)
Fine-Grained Privacy Extraction from Retrieval-Augmented Generation Systems via Knowledge Asymmetry Exploitation
por: Chen, Yufei, et al.
Publicado: (2025)
por: Chen, Yufei, et al.
Publicado: (2025)
TAPFixer: Automatic Detection and Repair of Home Automation Vulnerabilities based on Negated-property Reasoning
por: Yu, Yinbo, et al.
Publicado: (2024)
por: Yu, Yinbo, et al.
Publicado: (2024)
Can MLLMs Detect Phishing? A Comprehensive Security Benchmark Suite Focusing on Dynamic Threats and Multimodal Evaluation in Academic Environments
por: Zhou, Jingzhuo
Publicado: (2025)
por: Zhou, Jingzhuo
Publicado: (2025)
Black-Box Guardrail Reverse-engineering Attack
por: Yao, Hongwei, et al.
Publicado: (2025)
por: Yao, Hongwei, et al.
Publicado: (2025)
Slot: Provenance-Driven APT Detection through Graph Reinforcement Learning
por: Qiao, Wei, et al.
Publicado: (2024)
por: Qiao, Wei, et al.
Publicado: (2024)
SCAFFOLD-CEGIS: Preventing Latent Security Degradation in LLM-Driven Iterative Code Refinement
por: Chen, Yi, et al.
Publicado: (2026)
por: Chen, Yi, et al.
Publicado: (2026)
GhostEI-Bench: Do Mobile Agents Resilience to Environmental Injection in Dynamic On-Device Environments?
por: Chen, Chiyu, et al.
Publicado: (2025)
por: Chen, Chiyu, et al.
Publicado: (2025)
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms
por: He, Sinan, et al.
Publicado: (2025)
por: He, Sinan, et al.
Publicado: (2025)
On the Necessity of Pre-agreed Secrets for Thwarting Last-minute Coercion: Vulnerabilities and Lessons From the Loki E-voting Protocol
por: Qiao, Jingxin, et al.
Publicado: (2026)
por: Qiao, Jingxin, et al.
Publicado: (2026)
VeilAudit: Breaking the Deadlock Between Privacy and Accountability Across Blockchains
por: Qiao, Minhao, et al.
Publicado: (2025)
por: Qiao, Minhao, et al.
Publicado: (2025)
Building Intelligence Identification System via Large Language Model Watermarking: A Survey and Beyond
por: Wang, Xuhong, et al.
Publicado: (2024)
por: Wang, Xuhong, et al.
Publicado: (2024)
HarmChip: Evaluating Hardware Security Centric LLM Safety via Jailbreak Benchmarking
por: Wang, Zeng, et al.
Publicado: (2026)
por: Wang, Zeng, et al.
Publicado: (2026)
PraxiMLP: A Threshold-based Framework for Efficient Three-Party MLP with Practical Security
por: Tao, Tianle, et al.
Publicado: (2025)
por: Tao, Tianle, et al.
Publicado: (2025)
Behavioral Authentication for Security and Safety
por: Wang, Cheng, et al.
Publicado: (2023)
por: Wang, Cheng, et al.
Publicado: (2023)
Prompt Stealing Attacks Against Large Language Models
por: Sha, Zeyang, et al.
Publicado: (2024)
por: Sha, Zeyang, et al.
Publicado: (2024)
Demystifying Progressive Web Application Permission Systems
por: Wang, Mengxiao, et al.
Publicado: (2025)
por: Wang, Mengxiao, et al.
Publicado: (2025)
GuardianPWA: Enhancing Security Throughout the Progressive Web App Installation Lifecycle
por: Wang, Mengxiao, et al.
Publicado: (2025)
por: Wang, Mengxiao, et al.
Publicado: (2025)
MEOW: MEMOry Supervised LLM Unlearning Via Inverted Facts
por: Gu, Tianle, et al.
Publicado: (2024)
por: Gu, Tianle, et al.
Publicado: (2024)
Ejemplares similares
-
Probing the Robustness of Large Language Models Safety to Latent Perturbations
por: Gu, Tianle, et al.
Publicado: (2025) -
HoneypotNet: Backdoor Attacks Against Model Extraction
por: Wang, Yixu, et al.
Publicado: (2025) -
Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking
por: Gu, Tianle, et al.
Publicado: (2025) -
OpenRT: An Open-Source Red Teaming Framework for Multimodal LLMs
por: Wang, Xin, et al.
Publicado: (2026) -
StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic Data
por: Wang, Yixu, et al.
Publicado: (2025)