Assessment of Multimodal Large Language Models in Alignment with Human Values
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Zhelun, Wang, Zhipin, Fan, Hongxing, Zhang, Zaibin, Li, Lijun, Zhang, Yongting, Yin, Zhenfei, Sheng, Lu, Qiao, Yu, Shao, Jing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RH20T-P: A Primitive-Level Robotic Dataset Towards Composable Generalization Agents
von: Chen, Zeren, et al.
Veröffentlicht: (2024)
von: Chen, Zeren, et al.
Veröffentlicht: (2024)
SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model
von: Zhang, Yongting, et al.
Veröffentlicht: (2024)
von: Zhang, Yongting, et al.
Veröffentlicht: (2024)
PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety
von: Zhang, Zaibin, et al.
Veröffentlicht: (2024)
von: Zhang, Zaibin, et al.
Veröffentlicht: (2024)
From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities
von: Lu, Chaochao, et al.
Veröffentlicht: (2024)
von: Lu, Chaochao, et al.
Veröffentlicht: (2024)
WorldSimBench: Towards Video Generation Models as World Simulators
von: Qin, Yiran, et al.
Veröffentlicht: (2024)
von: Qin, Yiran, et al.
Veröffentlicht: (2024)
MineDreamer: Learning to Follow Instructions via Chain-of-Imagination for Simulated-World Control
von: Zhou, Enshen, et al.
Veröffentlicht: (2024)
von: Zhou, Enshen, et al.
Veröffentlicht: (2024)
MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active Perception
von: Qin, Yiran, et al.
Veröffentlicht: (2023)
von: Qin, Yiran, et al.
Veröffentlicht: (2023)
BEV-IO: Enhancing Bird's-Eye-View 3D Detection with Instance Occupancy
von: Zhang, Zaibin, et al.
Veröffentlicht: (2023)
von: Zhang, Zaibin, et al.
Veröffentlicht: (2023)
Systematic Reward Gap Optimization for Mitigating VLM Hallucinations
von: He, Lehan, et al.
Veröffentlicht: (2024)
von: He, Lehan, et al.
Veröffentlicht: (2024)
ProGuard: Towards Proactive Multimodal Safeguard
von: Yu, Shaohan, et al.
Veröffentlicht: (2025)
von: Yu, Shaohan, et al.
Veröffentlicht: (2025)
AD-H: Language-guided Autonomous Driving with Hierarchical Agents
von: Zhang, Zaibin, et al.
Veröffentlicht: (2024)
von: Zhang, Zaibin, et al.
Veröffentlicht: (2024)
Think3D: Thinking with Space for Spatial Reasoning
von: Zhang, Zaibin, et al.
Veröffentlicht: (2026)
von: Zhang, Zaibin, et al.
Veröffentlicht: (2026)
From Parts to Whole: A Unified Reference Framework for Controllable Human Image Generation
von: Huang, Zehuan, et al.
Veröffentlicht: (2024)
von: Huang, Zehuan, et al.
Veröffentlicht: (2024)
UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
von: Li, Teng, et al.
Veröffentlicht: (2025)
von: Li, Teng, et al.
Veröffentlicht: (2025)
Octavius: Mitigating Task Interference in MLLMs via LoRA-MoE
von: Chen, Zeren, et al.
Veröffentlicht: (2023)
von: Chen, Zeren, et al.
Veröffentlicht: (2023)
Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
EndoBench: A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy Analysis
von: Liu, Shengyuan, et al.
Veröffentlicht: (2025)
von: Liu, Shengyuan, et al.
Veröffentlicht: (2025)
InterMoE: Individual-Specific 3D Human Interaction Generation via Dynamic Temporal-Selective MoE
von: Wang, Lipeng, et al.
Veröffentlicht: (2025)
von: Wang, Lipeng, et al.
Veröffentlicht: (2025)
Reasoning-Driven Amodal Completion: Collaborative Agents and Perceptual Evaluation
von: Fan, Hongxing, et al.
Veröffentlicht: (2025)
von: Fan, Hongxing, et al.
Veröffentlicht: (2025)
Debiasing Multimodal Large Language Models via Penalization of Language Priors
von: Zhang, YiFan, et al.
Veröffentlicht: (2024)
von: Zhang, YiFan, et al.
Veröffentlicht: (2024)
GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation
von: Liang, Yupu, et al.
Veröffentlicht: (2025)
von: Liang, Yupu, et al.
Veröffentlicht: (2025)
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers
von: Qi, Daiqing, et al.
Veröffentlicht: (2025)
von: Qi, Daiqing, et al.
Veröffentlicht: (2025)
Depicting Beyond Scores: Advancing Image Quality Assessment through Multi-modal Language Models
von: You, Zhiyuan, et al.
Veröffentlicht: (2023)
von: You, Zhiyuan, et al.
Veröffentlicht: (2023)
Self-adaptive Dataset Construction for Real-World Multimodal Safety Scenarios
von: Qu, Jingen, et al.
Veröffentlicht: (2025)
von: Qu, Jingen, et al.
Veröffentlicht: (2025)
Ovis: Structural Embedding Alignment for Multimodal Large Language Model
von: Lu, Shiyin, et al.
Veröffentlicht: (2024)
von: Lu, Shiyin, et al.
Veröffentlicht: (2024)
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
von: Yan, Ziang, et al.
Veröffentlicht: (2024)
von: Yan, Ziang, et al.
Veröffentlicht: (2024)
ReDiPrune: Relevance-Diversity Pre-Projection Token Pruning for Efficient Multimodal LLMs
von: Yu, An, et al.
Veröffentlicht: (2026)
von: Yu, An, et al.
Veröffentlicht: (2026)
ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens
von: Lu, Zhuqiang, et al.
Veröffentlicht: (2024)
von: Lu, Zhuqiang, et al.
Veröffentlicht: (2024)
Semantic Alignment for Multimodal Large Language Models
von: Wu, Tao, et al.
Veröffentlicht: (2024)
von: Wu, Tao, et al.
Veröffentlicht: (2024)
TreeTeaming: Autonomous Red-Teaming of Vision-Language Models via Hierarchical Strategy Exploration
von: Li, Chunxiao, et al.
Veröffentlicht: (2026)
von: Li, Chunxiao, et al.
Veröffentlicht: (2026)
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
Towards Tracing Trustworthiness Dynamics: Revisiting Pre-training Period of Large Language Models
von: Qian, Chen, et al.
Veröffentlicht: (2024)
von: Qian, Chen, et al.
Veröffentlicht: (2024)
A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models
von: Li, Duo, et al.
Veröffentlicht: (2025)
von: Li, Duo, et al.
Veröffentlicht: (2025)
Gastric-X: A Multimodal Multi-Phase Benchmark Dataset for Advancing Vision-Language Models in Gastric Cancer Analysis
von: Lu, Sheng, et al.
Veröffentlicht: (2026)
von: Lu, Sheng, et al.
Veröffentlicht: (2026)
OLMD: Orientation-aware Long-term Motion Decoupling for Continuous Sign Language Recognition
von: Yu, Yiheng, et al.
Veröffentlicht: (2025)
von: Yu, Yiheng, et al.
Veröffentlicht: (2025)
Safety of Multimodal Large Language Models on Images and Texts
von: Liu, Xin, et al.
Veröffentlicht: (2024)
von: Liu, Xin, et al.
Veröffentlicht: (2024)
VIVA: A Benchmark for Vision-Grounded Decision-Making with Human Values
von: Hu, Zhe, et al.
Veröffentlicht: (2024)
von: Hu, Zhe, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
RH20T-P: A Primitive-Level Robotic Dataset Towards Composable Generalization Agents
von: Chen, Zeren, et al.
Veröffentlicht: (2024) -
SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model
von: Zhang, Yongting, et al.
Veröffentlicht: (2024) -
PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety
von: Zhang, Zaibin, et al.
Veröffentlicht: (2024) -
From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities
von: Lu, Chaochao, et al.
Veröffentlicht: (2024) -
WorldSimBench: Towards Video Generation Models as World Simulators
von: Qin, Yiran, et al.
Veröffentlicht: (2024)