MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Chejian, Zhang, Jiawei, Chen, Zhaorun, Xie, Chulin, Kang, Mintong, Potter, Yujin, Wang, Zhun, Yuan, Zhuowen, Xiong, Alexander, Xiong, Zidi, Zhang, Chenhui, Yuan, Lingzhi, Zeng, Yi, Xu, Peiyang, Guo, Chengquan, Zhou, Andy, Tan, Jeffrey Ziwei, Zhao, Xuandong, Pinto, Francesco, Xiang, Zhen, Gai, Yu, Lin, Zinan, Hendrycks, Dan, Li, Bo, Song, Dawn |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
di: Wang, Boxin, et al.
Pubblicazione: (2023)
di: Wang, Boxin, et al.
Pubblicazione: (2023)
Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset
di: Kang, Mintong, et al.
Pubblicazione: (2025)
di: Kang, Mintong, et al.
Pubblicazione: (2025)
AdvWave: Stealthy Adversarial Jailbreak Attack against Large Audio-Language Models
di: Kang, Mintong, et al.
Pubblicazione: (2024)
di: Kang, Mintong, et al.
Pubblicazione: (2024)
RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents
di: Guo, Chengquan, et al.
Pubblicazione: (2025)
di: Guo, Chengquan, et al.
Pubblicazione: (2025)
RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content
di: Yuan, Zhuowen, et al.
Pubblicazione: (2024)
di: Yuan, Zhuowen, et al.
Pubblicazione: (2024)
KnowHalu: Hallucination Detection via Multi-Form Knowledge Based Factual Checking
di: Zhang, Jiawei, et al.
Pubblicazione: (2024)
di: Zhang, Jiawei, et al.
Pubblicazione: (2024)
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
di: Hong, Junyuan, et al.
Pubblicazione: (2024)
di: Hong, Junyuan, et al.
Pubblicazione: (2024)
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
di: Guo, Chengquan, et al.
Pubblicazione: (2024)
di: Guo, Chengquan, et al.
Pubblicazione: (2024)
ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning
di: Chen, Zhaorun, et al.
Pubblicazione: (2025)
di: Chen, Zhaorun, et al.
Pubblicazione: (2025)
BlueCodeAgent: A Blue Teaming Agent Enabled by Automated Red Teaming for CodeGen AI
di: Guo, Chengquan, et al.
Pubblicazione: (2025)
di: Guo, Chengquan, et al.
Pubblicazione: (2025)
AdvAgent: Controllable Blackbox Red-teaming on Web Agents
di: Xu, Chejian, et al.
Pubblicazione: (2024)
di: Xu, Chejian, et al.
Pubblicazione: (2024)
VMDT: Decoding the Trustworthiness of Video Foundation Models
di: Potter, Yujin, et al.
Pubblicazione: (2025)
di: Potter, Yujin, et al.
Pubblicazione: (2025)
DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agents
di: Chen, Zhaorun, et al.
Pubblicazione: (2026)
di: Chen, Zhaorun, et al.
Pubblicazione: (2026)
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
di: Liao, Zeyi, et al.
Pubblicazione: (2024)
di: Liao, Zeyi, et al.
Pubblicazione: (2024)
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
di: Xiong, Alexander, et al.
Pubblicazione: (2025)
di: Xiong, Alexander, et al.
Pubblicazione: (2025)
ChatScene: Knowledge-Enabled Safety-Critical Scenario Generation for Autonomous Vehicles
di: Zhang, Jiawei, et al.
Pubblicazione: (2024)
di: Zhang, Jiawei, et al.
Pubblicazione: (2024)
DiffAttack: Evasion Attacks Against Diffusion-Based Adversarial Purification
di: Kang, Mintong, et al.
Pubblicazione: (2023)
di: Kang, Mintong, et al.
Pubblicazione: (2023)
GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning
di: Xiang, Zhen, et al.
Pubblicazione: (2024)
di: Xiang, Zhen, et al.
Pubblicazione: (2024)
CBD: A Certified Backdoor Detector Based on Local Dominant Probability
di: Xiang, Zhen, et al.
Pubblicazione: (2023)
di: Xiang, Zhen, et al.
Pubblicazione: (2023)
Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
di: Xiong, Zidi, et al.
Pubblicazione: (2026)
di: Xiong, Zidi, et al.
Pubblicazione: (2026)
AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents
di: Xie, Jingxu, et al.
Pubblicazione: (2025)
di: Xie, Jingxu, et al.
Pubblicazione: (2025)
Frontier AI's Impact on the Cybersecurity Landscape
di: Potter, Yujin, et al.
Pubblicazione: (2025)
di: Potter, Yujin, et al.
Pubblicazione: (2025)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
di: Yuan, Lingzhi, et al.
Pubblicazione: (2025)
di: Yuan, Lingzhi, et al.
Pubblicazione: (2025)
SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability
di: Xu, Peiyang, et al.
Pubblicazione: (2025)
di: Xu, Peiyang, et al.
Pubblicazione: (2025)
Message Passing Based Demodulation of the Time-Encoded Digital Modulation Signal
di: Xu, Yuan, et al.
Pubblicazione: (2025)
di: Xu, Yuan, et al.
Pubblicazione: (2025)
The Effects of Anlotinib Combined with Chemotherapy following Progression on Cyclin‐Dependent Kinase 4/6 Inhibitor in Hormone Receptor‐Positive Metastatic Breast Cancer
di: Ting Xu, et al.
Pubblicazione: (2024)
di: Ting Xu, et al.
Pubblicazione: (2024)
LLM-PBE: Assessing Data Privacy in Large Language Models
di: Li, Qinbin, et al.
Pubblicazione: (2024)
di: Li, Qinbin, et al.
Pubblicazione: (2024)
ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play Attacks
di: Chen, Zhaorun, et al.
Pubblicazione: (2025)
di: Chen, Zhaorun, et al.
Pubblicazione: (2025)
Soft Tail-dropping for Adaptive Visual Tokenization
di: Chen, Zeyuan, et al.
Pubblicazione: (2026)
di: Chen, Zeyuan, et al.
Pubblicazione: (2026)
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
di: Nie, Yuzhou, et al.
Pubblicazione: (2024)
di: Nie, Yuzhou, et al.
Pubblicazione: (2024)
Self-Sovereign Agent
di: Qu, Wenjie, et al.
Pubblicazione: (2026)
di: Qu, Wenjie, et al.
Pubblicazione: (2026)
Hidden Persuaders: LLMs' Political Leaning and Their Influence on Voters
di: Potter, Yujin, et al.
Pubblicazione: (2024)
di: Potter, Yujin, et al.
Pubblicazione: (2024)
Peer-Preservation in Frontier Models
di: Potter, Yujin, et al.
Pubblicazione: (2026)
di: Potter, Yujin, et al.
Pubblicazione: (2026)
An Undetectable Watermark for Generative Image Models
di: Gunn, Sam, et al.
Pubblicazione: (2024)
di: Gunn, Sam, et al.
Pubblicazione: (2024)
Scalable Best-of-N Selection for Large Language Models via Self-Certainty
di: Kang, Zhewei, et al.
Pubblicazione: (2025)
di: Kang, Zhewei, et al.
Pubblicazione: (2025)
VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection
di: Nie, Yuzhou, et al.
Pubblicazione: (2025)
di: Nie, Yuzhou, et al.
Pubblicazione: (2025)
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models
di: Xiong, Zidi, et al.
Pubblicazione: (2025)
di: Xiong, Zidi, et al.
Pubblicazione: (2025)
Efficacy and safety of nanoparticle albumin‐bound paclitaxel in taxane‐pretreated metastatic breast cancer patients
di: Weili Xiong, et al.
Pubblicazione: (2024)
di: Weili Xiong, et al.
Pubblicazione: (2024)
User-Assistant Bias in LLMs
di: Pan, Xu, et al.
Pubblicazione: (2025)
di: Pan, Xu, et al.
Pubblicazione: (2025)
Chain-of-Adaptation: Surgical Vision-Language Adaptation with Reinforcement Learning
di: Li, Jiajie, et al.
Pubblicazione: (2026)
di: Li, Jiajie, et al.
Pubblicazione: (2026)
Documenti analoghi
-
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
di: Wang, Boxin, et al.
Pubblicazione: (2023) -
Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset
di: Kang, Mintong, et al.
Pubblicazione: (2025) -
AdvWave: Stealthy Adversarial Jailbreak Attack against Large Audio-Language Models
di: Kang, Mintong, et al.
Pubblicazione: (2024) -
RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents
di: Guo, Chengquan, et al.
Pubblicazione: (2025) -
RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content
di: Yuan, Zhuowen, et al.
Pubblicazione: (2024)