From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Shixin, Liang, Jiafeng, Wang, Jiyuan, Dong, Xuan, Chang, Heng, Yu, Weijiang, Du, Jinhua, Liu, Ming, Qin, Bing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Affordance Benchmark for MLLMs
von: Wang, Junying, et al.
Veröffentlicht: (2025)
von: Wang, Junying, et al.
Veröffentlicht: (2025)
Linking Perception, Confidence and Accuracy in MLLMs
von: Du, Yuetian, et al.
Veröffentlicht: (2026)
von: Du, Yuetian, et al.
Veröffentlicht: (2026)
MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages
von: Du, Yexing, et al.
Veröffentlicht: (2025)
von: Du, Yexing, et al.
Veröffentlicht: (2025)
ChartEdit: How Far Are MLLMs From Automating Chart Analysis? Evaluating MLLMs' Capability via Chart Editing
von: Zhao, Xuanle, et al.
Veröffentlicht: (2025)
von: Zhao, Xuanle, et al.
Veröffentlicht: (2025)
Redundancy Principles for MLLMs Benchmarks
von: Zhang, Zicheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zicheng, et al.
Veröffentlicht: (2025)
AM$^3$Safety: Towards Data Efficient Alignment of Multi-modal Multi-turn Safety for MLLMs
von: Zhu, Han, et al.
Veröffentlicht: (2026)
von: Zhu, Han, et al.
Veröffentlicht: (2026)
Do MLLMs Really Understand the Charts?
von: Zhang, Xiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xiao, et al.
Veröffentlicht: (2025)
RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees
von: Xu, Yichen, et al.
Veröffentlicht: (2026)
von: Xu, Yichen, et al.
Veröffentlicht: (2026)
From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
Unhackable Temporal Rewarding for Scalable Video MLLMs
von: Yu, En, et al.
Veröffentlicht: (2025)
von: Yu, En, et al.
Veröffentlicht: (2025)
LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops
von: Fu, Jiyuan, et al.
Veröffentlicht: (2025)
von: Fu, Jiyuan, et al.
Veröffentlicht: (2025)
OCR or Not? Rethinking Document Information Extraction in the MLLMs Era with Real-World Large-Scale Datasets
von: Shen, Jiyuan, et al.
Veröffentlicht: (2026)
von: Shen, Jiyuan, et al.
Veröffentlicht: (2026)
OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2025)
Roles of MLLMs in Visually Rich Document Retrieval for RAG: A Survey
von: Zhang, Xiantao
Veröffentlicht: (2025)
von: Zhang, Xiantao
Veröffentlicht: (2025)
AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
von: Wang, Junyang, et al.
Veröffentlicht: (2023)
von: Wang, Junyang, et al.
Veröffentlicht: (2023)
GRIT: Teaching MLLMs to Think with Images
von: Fan, Yue, et al.
Veröffentlicht: (2025)
von: Fan, Yue, et al.
Veröffentlicht: (2025)
The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image Reasoning
von: Chen, Renmiao, et al.
Veröffentlicht: (2026)
von: Chen, Renmiao, et al.
Veröffentlicht: (2026)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs
von: Ye, Hengwei, et al.
Veröffentlicht: (2026)
von: Ye, Hengwei, et al.
Veröffentlicht: (2026)
When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
von: Liu, Hongcheng, et al.
Veröffentlicht: (2025)
von: Liu, Hongcheng, et al.
Veröffentlicht: (2025)
MileBench: Benchmarking MLLMs in Long Context
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
MM-THEBench: Do Reasoning MLLMs Think Reasonably?
von: Huang, Zhidian, et al.
Veröffentlicht: (2026)
von: Huang, Zhidian, et al.
Veröffentlicht: (2026)
Visual Room 2.0: Seeing is Not Understanding for MLLMs
von: Li, Haokun, et al.
Veröffentlicht: (2025)
von: Li, Haokun, et al.
Veröffentlicht: (2025)
MLLMs-Augmented Visual-Language Representation Learning
von: Liu, Yanqing, et al.
Veröffentlicht: (2023)
von: Liu, Yanqing, et al.
Veröffentlicht: (2023)
AdaCodec: A Predictive Visual Code for Video MLLMs
von: Hou, Haowen, et al.
Veröffentlicht: (2026)
von: Hou, Haowen, et al.
Veröffentlicht: (2026)
Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations
von: Li, Shuo, et al.
Veröffentlicht: (2025)
von: Li, Shuo, et al.
Veröffentlicht: (2025)
What MLLMs Learn about When they Learn about Multimodal Reasoning
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
Exploring the Design Space of Visual Context Representation in Video MLLMs
von: Du, Yifan, et al.
Veröffentlicht: (2024)
von: Du, Yifan, et al.
Veröffentlicht: (2024)
COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMs
von: Guo, Peizheng, et al.
Veröffentlicht: (2025)
von: Guo, Peizheng, et al.
Veröffentlicht: (2025)
Can MLLMs Perform Text-to-Image In-Context Learning?
von: Zeng, Yuchen, et al.
Veröffentlicht: (2024)
von: Zeng, Yuchen, et al.
Veröffentlicht: (2024)
Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification?
von: Lin, Fei, et al.
Veröffentlicht: (2025)
von: Lin, Fei, et al.
Veröffentlicht: (2025)
Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
von: Zhang, Huanyu, et al.
Veröffentlicht: (2025)
von: Zhang, Huanyu, et al.
Veröffentlicht: (2025)
GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning
von: Chen, Guizhen, et al.
Veröffentlicht: (2025)
von: Chen, Guizhen, et al.
Veröffentlicht: (2025)
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
von: Yeh, Chun-Hsiao, et al.
Veröffentlicht: (2025)
von: Yeh, Chun-Hsiao, et al.
Veröffentlicht: (2025)
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs
von: Ding, Xuanwen, et al.
Veröffentlicht: (2025)
von: Ding, Xuanwen, et al.
Veröffentlicht: (2025)
Distill Visual Chart Reasoning Ability from LLMs to MLLMs
von: He, Wei, et al.
Veröffentlicht: (2024)
von: He, Wei, et al.
Veröffentlicht: (2024)
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
The Forgotten Shield: Safety Grafting in Parameter-Space for Medical MLLMs
von: Zhao, Jiale, et al.
Veröffentlicht: (2025)
von: Zhao, Jiale, et al.
Veröffentlicht: (2025)
Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression
von: Du, Yao, et al.
Veröffentlicht: (2026)
von: Du, Yao, et al.
Veröffentlicht: (2026)
FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow
von: Sun, Haoyu, et al.
Veröffentlicht: (2025)
von: Sun, Haoyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Affordance Benchmark for MLLMs
von: Wang, Junying, et al.
Veröffentlicht: (2025) -
Linking Perception, Confidence and Accuracy in MLLMs
von: Du, Yuetian, et al.
Veröffentlicht: (2026) -
MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages
von: Du, Yexing, et al.
Veröffentlicht: (2025) -
ChartEdit: How Far Are MLLMs From Automating Chart Analysis? Evaluating MLLMs' Capability via Chart Editing
von: Zhao, Xuanle, et al.
Veröffentlicht: (2025) -
Redundancy Principles for MLLMs Benchmarks
von: Zhang, Zicheng, et al.
Veröffentlicht: (2025)