From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Jiang, Shixin, Liang, Jiafeng, Wang, Jiyuan, Dong, Xuan, Chang, Heng, Yu, Weijiang, Du, Jinhua, Liu, Ming, Qin, Bing |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Affordance Benchmark for MLLMs
par: Wang, Junying, et autres
Publié: (2025)
par: Wang, Junying, et autres
Publié: (2025)
Linking Perception, Confidence and Accuracy in MLLMs
par: Du, Yuetian, et autres
Publié: (2026)
par: Du, Yuetian, et autres
Publié: (2026)
MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages
par: Du, Yexing, et autres
Publié: (2025)
par: Du, Yexing, et autres
Publié: (2025)
ChartEdit: How Far Are MLLMs From Automating Chart Analysis? Evaluating MLLMs' Capability via Chart Editing
par: Zhao, Xuanle, et autres
Publié: (2025)
par: Zhao, Xuanle, et autres
Publié: (2025)
Redundancy Principles for MLLMs Benchmarks
par: Zhang, Zicheng, et autres
Publié: (2025)
par: Zhang, Zicheng, et autres
Publié: (2025)
AM$^3$Safety: Towards Data Efficient Alignment of Multi-modal Multi-turn Safety for MLLMs
par: Zhu, Han, et autres
Publié: (2026)
par: Zhu, Han, et autres
Publié: (2026)
Do MLLMs Really Understand the Charts?
par: Zhang, Xiao, et autres
Publié: (2025)
par: Zhang, Xiao, et autres
Publié: (2025)
RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees
par: Xu, Yichen, et autres
Publié: (2026)
par: Xu, Yichen, et autres
Publié: (2026)
From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking
par: Wang, Siyuan, et autres
Publié: (2024)
par: Wang, Siyuan, et autres
Publié: (2024)
Unhackable Temporal Rewarding for Scalable Video MLLMs
par: Yu, En, et autres
Publié: (2025)
par: Yu, En, et autres
Publié: (2025)
LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops
par: Fu, Jiyuan, et autres
Publié: (2025)
par: Fu, Jiyuan, et autres
Publié: (2025)
OCR or Not? Rethinking Document Information Extraction in the MLLMs Era with Real-World Large-Scale Datasets
par: Shen, Jiyuan, et autres
Publié: (2026)
par: Shen, Jiyuan, et autres
Publié: (2026)
OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
par: Zhao, Xiangyu, et autres
Publié: (2025)
par: Zhao, Xiangyu, et autres
Publié: (2025)
Roles of MLLMs in Visually Rich Document Retrieval for RAG: A Survey
par: Zhang, Xiantao
Publié: (2025)
par: Zhang, Xiantao
Publié: (2025)
AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
par: Wang, Junyang, et autres
Publié: (2023)
par: Wang, Junyang, et autres
Publié: (2023)
GRIT: Teaching MLLMs to Think with Images
par: Fan, Yue, et autres
Publié: (2025)
par: Fan, Yue, et autres
Publié: (2025)
The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image Reasoning
par: Chen, Renmiao, et autres
Publié: (2026)
par: Chen, Renmiao, et autres
Publié: (2026)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
par: Lu, Yujie, et autres
Publié: (2024)
par: Lu, Yujie, et autres
Publié: (2024)
Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs
par: Ye, Hengwei, et autres
Publié: (2026)
par: Ye, Hengwei, et autres
Publié: (2026)
When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
par: Liu, Hongcheng, et autres
Publié: (2025)
par: Liu, Hongcheng, et autres
Publié: (2025)
MileBench: Benchmarking MLLMs in Long Context
par: Song, Dingjie, et autres
Publié: (2024)
par: Song, Dingjie, et autres
Publié: (2024)
MM-THEBench: Do Reasoning MLLMs Think Reasonably?
par: Huang, Zhidian, et autres
Publié: (2026)
par: Huang, Zhidian, et autres
Publié: (2026)
Visual Room 2.0: Seeing is Not Understanding for MLLMs
par: Li, Haokun, et autres
Publié: (2025)
par: Li, Haokun, et autres
Publié: (2025)
MLLMs-Augmented Visual-Language Representation Learning
par: Liu, Yanqing, et autres
Publié: (2023)
par: Liu, Yanqing, et autres
Publié: (2023)
AdaCodec: A Predictive Visual Code for Video MLLMs
par: Hou, Haowen, et autres
Publié: (2026)
par: Hou, Haowen, et autres
Publié: (2026)
Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations
par: Li, Shuo, et autres
Publié: (2025)
par: Li, Shuo, et autres
Publié: (2025)
What MLLMs Learn about When they Learn about Multimodal Reasoning
par: Chung, Jiwan, et autres
Publié: (2025)
par: Chung, Jiwan, et autres
Publié: (2025)
Exploring the Design Space of Visual Context Representation in Video MLLMs
par: Du, Yifan, et autres
Publié: (2024)
par: Du, Yifan, et autres
Publié: (2024)
COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMs
par: Guo, Peizheng, et autres
Publié: (2025)
par: Guo, Peizheng, et autres
Publié: (2025)
Can MLLMs Perform Text-to-Image In-Context Learning?
par: Zeng, Yuchen, et autres
Publié: (2024)
par: Zeng, Yuchen, et autres
Publié: (2024)
Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification?
par: Lin, Fei, et autres
Publié: (2025)
par: Lin, Fei, et autres
Publié: (2025)
Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
par: Zhang, Huanyu, et autres
Publié: (2025)
par: Zhang, Huanyu, et autres
Publié: (2025)
GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning
par: Chen, Guizhen, et autres
Publié: (2025)
par: Chen, Guizhen, et autres
Publié: (2025)
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
par: Yeh, Chun-Hsiao, et autres
Publié: (2025)
par: Yeh, Chun-Hsiao, et autres
Publié: (2025)
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs
par: Ding, Xuanwen, et autres
Publié: (2025)
par: Ding, Xuanwen, et autres
Publié: (2025)
Distill Visual Chart Reasoning Ability from LLMs to MLLMs
par: He, Wei, et autres
Publié: (2024)
par: He, Wei, et autres
Publié: (2024)
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
par: Wang, Yuhao, et autres
Publié: (2024)
par: Wang, Yuhao, et autres
Publié: (2024)
The Forgotten Shield: Safety Grafting in Parameter-Space for Medical MLLMs
par: Zhao, Jiale, et autres
Publié: (2025)
par: Zhao, Jiale, et autres
Publié: (2025)
Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression
par: Du, Yao, et autres
Publié: (2026)
par: Du, Yao, et autres
Publié: (2026)
FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow
par: Sun, Haoyu, et autres
Publié: (2025)
par: Sun, Haoyu, et autres
Publié: (2025)
Documents similaires
-
Affordance Benchmark for MLLMs
par: Wang, Junying, et autres
Publié: (2025) -
Linking Perception, Confidence and Accuracy in MLLMs
par: Du, Yuetian, et autres
Publié: (2026) -
MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages
par: Du, Yexing, et autres
Publié: (2025) -
ChartEdit: How Far Are MLLMs From Automating Chart Analysis? Evaluating MLLMs' Capability via Chart Editing
par: Zhao, Xuanle, et autres
Publié: (2025) -
Redundancy Principles for MLLMs Benchmarks
par: Zhang, Zicheng, et autres
Publié: (2025)