Matryoshka Multimodal Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cai, Mu, Yang, Jianwei, Gao, Jianfeng, Lee, Yong Jae |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
von: Cai, Mu, et al.
Veröffentlicht: (2024)
von: Cai, Mu, et al.
Veröffentlicht: (2024)
Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos
von: Zhang, Jianrui, et al.
Veröffentlicht: (2024)
von: Zhang, Jianrui, et al.
Veröffentlicht: (2024)
CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples
von: Zhang, Jianrui, et al.
Veröffentlicht: (2024)
von: Zhang, Jianrui, et al.
Veröffentlicht: (2024)
ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
von: Cai, Mu, et al.
Veröffentlicht: (2023)
von: Cai, Mu, et al.
Veröffentlicht: (2023)
Leveraging Large Language Models for Scalable Vector Graphics-Driven Image Understanding
von: Cai, Mu, et al.
Veröffentlicht: (2023)
von: Cai, Mu, et al.
Veröffentlicht: (2023)
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
von: Shang, Yuzhang, et al.
Veröffentlicht: (2024)
von: Shang, Yuzhang, et al.
Veröffentlicht: (2024)
VGBench: Evaluating Large Language Models on Vector Graphics Understanding and Generation
von: Zou, Bocheng, et al.
Veröffentlicht: (2024)
von: Zou, Bocheng, et al.
Veröffentlicht: (2024)
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
von: Yu, Weihao, et al.
Veröffentlicht: (2023)
von: Yu, Weihao, et al.
Veröffentlicht: (2023)
Improved Baselines with Visual Instruction Tuning
von: Liu, Haotian, et al.
Veröffentlicht: (2023)
von: Liu, Haotian, et al.
Veröffentlicht: (2023)
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
By My Eyes: Grounding Multimodal Large Language Models with Sensor Data via Visual Prompting
von: Yoon, Hyungjun, et al.
Veröffentlicht: (2024)
von: Yoon, Hyungjun, et al.
Veröffentlicht: (2024)
OmniParser for Pure Vision Based GUI Agent
von: Lu, Yadong, et al.
Veröffentlicht: (2024)
von: Lu, Yadong, et al.
Veröffentlicht: (2024)
MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning
von: Cai, Zikui, et al.
Veröffentlicht: (2025)
von: Cai, Zikui, et al.
Veröffentlicht: (2025)
Towards Understanding Graphical Perception in Large Multimodal Models
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
Magma: A Foundation Model for Multimodal AI Agents
von: Yang, Jianwei, et al.
Veröffentlicht: (2025)
von: Yang, Jianwei, et al.
Veröffentlicht: (2025)
A Concept-based Interpretable Model for the Diagnosis of Choroid Neoplasias using Multimodal Data
von: Wu, Yifan, et al.
Veröffentlicht: (2024)
von: Wu, Yifan, et al.
Veröffentlicht: (2024)
Benchmarking the Thinking Mode of Multimodal Large Language Models in Clinical Tasks
von: Hong, Jindong, et al.
Veröffentlicht: (2025)
von: Hong, Jindong, et al.
Veröffentlicht: (2025)
Real Deep Research for AI, Robotics and Beyond
von: Zou, Xueyan, et al.
Veröffentlicht: (2025)
von: Zou, Xueyan, et al.
Veröffentlicht: (2025)
Ovis: Structural Embedding Alignment for Multimodal Large Language Model
von: Lu, Shiyin, et al.
Veröffentlicht: (2024)
von: Lu, Shiyin, et al.
Veröffentlicht: (2024)
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
von: Yu, Zhuoran, et al.
Veröffentlicht: (2025)
von: Yu, Zhuoran, et al.
Veröffentlicht: (2025)
Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection
von: Mei, Jingbiao, et al.
Veröffentlicht: (2025)
von: Mei, Jingbiao, et al.
Veröffentlicht: (2025)
MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
von: Lu, Pan, et al.
Veröffentlicht: (2023)
von: Lu, Pan, et al.
Veröffentlicht: (2023)
List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs
von: Yan, An, et al.
Veröffentlicht: (2024)
von: Yan, An, et al.
Veröffentlicht: (2024)
Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
von: Chen, Junying, et al.
Veröffentlicht: (2024)
von: Chen, Junying, et al.
Veröffentlicht: (2024)
Bridging the Gap Between Multimodal Foundation Models and World Models
von: He, Xuehai
Veröffentlicht: (2025)
von: He, Xuehai
Veröffentlicht: (2025)
Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models
von: Wang, Hengyi, et al.
Veröffentlicht: (2024)
von: Wang, Hengyi, et al.
Veröffentlicht: (2024)
HEMM: Holistic Evaluation of Multimodal Foundation Models
von: Liang, Paul Pu, et al.
Veröffentlicht: (2024)
von: Liang, Paul Pu, et al.
Veröffentlicht: (2024)
A Survey on Multimodal Large Language Models
von: Yin, Shukang, et al.
Veröffentlicht: (2023)
von: Yin, Shukang, et al.
Veröffentlicht: (2023)
Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging
von: Cai, Zhenyang, et al.
Veröffentlicht: (2024)
von: Cai, Zhenyang, et al.
Veröffentlicht: (2024)
Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models
von: Zeng, Yu, et al.
Veröffentlicht: (2026)
von: Zeng, Yu, et al.
Veröffentlicht: (2026)
OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents
von: Yang, Rui, et al.
Veröffentlicht: (2026)
von: Yang, Rui, et al.
Veröffentlicht: (2026)
Many-Shot In-Context Learning in Multimodal Foundation Models
von: Jiang, Yixing, et al.
Veröffentlicht: (2024)
von: Jiang, Yixing, et al.
Veröffentlicht: (2024)
Visual Question Decomposition on Multimodal Large Language Models
von: Zhang, Haowei, et al.
Veröffentlicht: (2024)
von: Zhang, Haowei, et al.
Veröffentlicht: (2024)
LMFusion: Adapting Pretrained Language Models for Multimodal Generation
von: Shi, Weijia, et al.
Veröffentlicht: (2024)
von: Shi, Weijia, et al.
Veröffentlicht: (2024)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
Woodpecker: Hallucination Correction for Multimodal Large Language Models
von: Yin, Shukang, et al.
Veröffentlicht: (2023)
von: Yin, Shukang, et al.
Veröffentlicht: (2023)
Promptception: How Sensitive Are Large Multimodal Models to Prompts?
von: Ismithdeen, Mohamed Insaf, et al.
Veröffentlicht: (2025)
von: Ismithdeen, Mohamed Insaf, et al.
Veröffentlicht: (2025)
Matryoshka Query Transformer for Large Vision-Language Models
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
MatMamba: A Matryoshka State Space Model
von: Shukla, Abhinav, et al.
Veröffentlicht: (2024)
von: Shukla, Abhinav, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
von: Cai, Mu, et al.
Veröffentlicht: (2024) -
Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos
von: Zhang, Jianrui, et al.
Veröffentlicht: (2024) -
CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples
von: Zhang, Jianrui, et al.
Veröffentlicht: (2024) -
ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
von: Cai, Mu, et al.
Veröffentlicht: (2023) -
Leveraging Large Language Models for Scalable Vector Graphics-Driven Image Understanding
von: Cai, Mu, et al.
Veröffentlicht: (2023)