VMDT: Decoding the Trustworthiness of Video Foundation Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Potter, Yujin, Wang, Zhun, Crispino, Nicholas, Montgomery, Kyle, Xiong, Alexander, Chang, Ethan Y., Pinto, Francesco, Chen, Yuqi, Gupta, Rahul, Ziyadi, Morteza, Christodoulopoulos, Christos, Li, Bo, Wang, Chenguang, Song, Dawn |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Peer-Preservation in Frontier Models
di: Potter, Yujin, et al.
Pubblicazione: (2026)
di: Potter, Yujin, et al.
Pubblicazione: (2026)
Agent Instructs Large Language Models to be General Zero-Shot Reasoners
di: Crispino, Nicholas, et al.
Pubblicazione: (2023)
di: Crispino, Nicholas, et al.
Pubblicazione: (2023)
MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models
di: Xu, Chejian, et al.
Pubblicazione: (2025)
di: Xu, Chejian, et al.
Pubblicazione: (2025)
A Framework for Formalizing LLM Agent Security
di: Siu, Vincent, et al.
Pubblicazione: (2026)
di: Siu, Vincent, et al.
Pubblicazione: (2026)
COSMIC: Generalized Refusal Direction Identification in LLM Activations
di: Siu, Vincent, et al.
Pubblicazione: (2025)
di: Siu, Vincent, et al.
Pubblicazione: (2025)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
di: Siu, Vincent, et al.
Pubblicazione: (2025)
di: Siu, Vincent, et al.
Pubblicazione: (2025)
Re-Tuning: Overcoming the Compositionality Limits of Large Language Models with Recursive Tuning
di: Pasewark, Eric, et al.
Pubblicazione: (2024)
di: Pasewark, Eric, et al.
Pubblicazione: (2024)
LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
di: Kolasani, Sai, et al.
Pubblicazione: (2025)
di: Kolasani, Sai, et al.
Pubblicazione: (2025)
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
di: Siu, Vincent, et al.
Pubblicazione: (2025)
di: Siu, Vincent, et al.
Pubblicazione: (2025)
Certifying Counterfactual Bias in LLMs
di: Chaudhary, Isha, et al.
Pubblicazione: (2024)
di: Chaudhary, Isha, et al.
Pubblicazione: (2024)
Predicting Task Performance with Context-aware Scaling Laws
di: Montgomery, Kyle, et al.
Pubblicazione: (2025)
di: Montgomery, Kyle, et al.
Pubblicazione: (2025)
Customize Multi-modal RAI Guardrails with Precedent-based predictions
di: Yang, Cheng-Fu, et al.
Pubblicazione: (2025)
di: Yang, Cheng-Fu, et al.
Pubblicazione: (2025)
Benchmarking Zero-Shot Robustness of Multimodal Foundation Models: A Pilot Study
di: Wang, Chenguang, et al.
Pubblicazione: (2024)
di: Wang, Chenguang, et al.
Pubblicazione: (2024)
Budget-aware Test-time Scaling via Discriminative Verification
di: Montgomery, Kyle, et al.
Pubblicazione: (2025)
di: Montgomery, Kyle, et al.
Pubblicazione: (2025)
Partial Federated Learning
di: Feng, Tiantian, et al.
Pubblicazione: (2024)
di: Feng, Tiantian, et al.
Pubblicazione: (2024)
Frontier AI's Impact on the Cybersecurity Landscape
di: Potter, Yujin, et al.
Pubblicazione: (2025)
di: Potter, Yujin, et al.
Pubblicazione: (2025)
Aligning AI with Public Values: Deliberation and Decision-Making for Governing Multimodal LLMs in Political Video Analysis
di: Sharma, Tanusree, et al.
Pubblicazione: (2024)
di: Sharma, Tanusree, et al.
Pubblicazione: (2024)
InterPrior: Scaling Generative Control for Physics-Based Human-Object Interactions
di: Xu, Sirui, et al.
Pubblicazione: (2026)
di: Xu, Sirui, et al.
Pubblicazione: (2026)
Hidden Persuaders: LLMs' Political Leaning and Their Influence on Voters
di: Potter, Yujin, et al.
Pubblicazione: (2024)
di: Potter, Yujin, et al.
Pubblicazione: (2024)
Future of Algorithmic Organization: Large-Scale Analysis of Decentralized Autonomous Organizations (DAOs)
di: Sharma, Tanusree, et al.
Pubblicazione: (2024)
di: Sharma, Tanusree, et al.
Pubblicazione: (2024)
SoK: The Gap Between Data Rights Ideals and Reality
di: Potter, Yujin, et al.
Pubblicazione: (2023)
di: Potter, Yujin, et al.
Pubblicazione: (2023)
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
di: Wang, Boxin, et al.
Pubblicazione: (2023)
di: Wang, Boxin, et al.
Pubblicazione: (2023)
MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models
di: Tu, Jianhong, et al.
Pubblicazione: (2024)
di: Tu, Jianhong, et al.
Pubblicazione: (2024)
C-MORAL: Controllable Multi-Objective Molecular Optimization with Reinforcement Alignment for LLMs
di: Gao, Rui, et al.
Pubblicazione: (2026)
di: Gao, Rui, et al.
Pubblicazione: (2026)
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
di: Wang, Zhun, et al.
Pubblicazione: (2025)
di: Wang, Zhun, et al.
Pubblicazione: (2025)
Evolving AI Collectives to Enhance Human Diversity and Enable Self-Regulation
di: Lai, Shiyang, et al.
Pubblicazione: (2024)
di: Lai, Shiyang, et al.
Pubblicazione: (2024)
OPTAGENT: Optimizing Multi-Agent LLM Interactions Through Verbal Reinforcement Learning for Enhanced Reasoning
di: Bi, Zhenyu, et al.
Pubblicazione: (2025)
di: Bi, Zhenyu, et al.
Pubblicazione: (2025)
JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning Evaluation
di: Bi, Zhenyu, et al.
Pubblicazione: (2025)
di: Bi, Zhenyu, et al.
Pubblicazione: (2025)
SWAN: Semantic Watermarking with Abstract Meaning Representation
di: Ye, Ziping, et al.
Pubblicazione: (2026)
di: Ye, Ziping, et al.
Pubblicazione: (2026)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
di: Wang, Xilong, et al.
Pubblicazione: (2026)
di: Wang, Xilong, et al.
Pubblicazione: (2026)
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
di: Hong, Junyuan, et al.
Pubblicazione: (2024)
di: Hong, Junyuan, et al.
Pubblicazione: (2024)
PRINCIPIOS DE NEUROBIOLOGÍA
di: Luis Crispino
Pubblicazione: (2011)
di: Luis Crispino
Pubblicazione: (2011)
BeyondBench: Contamination-Resistant Evaluation of Reasoning in Language Models
di: Srivastava, Gaurav, et al.
Pubblicazione: (2025)
di: Srivastava, Gaurav, et al.
Pubblicazione: (2025)
RARE: Retrieval-Aware Robustness Evaluation for Retrieval-Augmented Generation Systems
di: Zeng, Yixiao, et al.
Pubblicazione: (2025)
di: Zeng, Yixiao, et al.
Pubblicazione: (2025)
The Media Specialist...Between A Rock and A Hard Place.
di: Heller, Dawn, Ed., et al.
Pubblicazione: (1978)
di: Heller, Dawn, Ed., et al.
Pubblicazione: (1978)
Harmonic LLMs are Trustworthy
di: Kersting, Nicholas S., et al.
Pubblicazione: (2024)
di: Kersting, Nicholas S., et al.
Pubblicazione: (2024)
RAGPPI: RAG Benchmark for Protein-Protein Interactions in Drug Discovery
di: Jeon, Youngseung, et al.
Pubblicazione: (2025)
di: Jeon, Youngseung, et al.
Pubblicazione: (2025)
Probabilistic Conceptual Explainers: Trustworthy Conceptual Explanations for Vision Foundation Models
di: Wang, Hengyi, et al.
Pubblicazione: (2024)
di: Wang, Hengyi, et al.
Pubblicazione: (2024)
On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective
di: Huang, Yue, et al.
Pubblicazione: (2025)
di: Huang, Yue, et al.
Pubblicazione: (2025)
On the Foundations of Trustworthy Artificial Intelligence
di: Dunham, TJ
Pubblicazione: (2026)
di: Dunham, TJ
Pubblicazione: (2026)
Documenti analoghi
-
Peer-Preservation in Frontier Models
di: Potter, Yujin, et al.
Pubblicazione: (2026) -
Agent Instructs Large Language Models to be General Zero-Shot Reasoners
di: Crispino, Nicholas, et al.
Pubblicazione: (2023) -
MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models
di: Xu, Chejian, et al.
Pubblicazione: (2025) -
A Framework for Formalizing LLM Agent Security
di: Siu, Vincent, et al.
Pubblicazione: (2026) -
COSMIC: Generalized Refusal Direction Identification in LLM Activations
di: Siu, Vincent, et al.
Pubblicazione: (2025)