Founder effects shape the evolutionary dynamics of multimodality in open LLM families
Fuente:
arXiv
Salvato in:
| Autore principale: | Cebrian, Manuel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Automatic benchmarking of large multimodal models via iterative experiment programming
di: Conti, Alessandro, et al.
Pubblicazione: (2024)
di: Conti, Alessandro, et al.
Pubblicazione: (2024)
GlitchBench: Can large multimodal models detect video game glitches?
di: Taesiri, Mohammad Reza, et al.
Pubblicazione: (2023)
di: Taesiri, Mohammad Reza, et al.
Pubblicazione: (2023)
MAIRA-1: A specialised large multimodal model for radiology report generation
di: Hyland, Stephanie L., et al.
Pubblicazione: (2023)
di: Hyland, Stephanie L., et al.
Pubblicazione: (2023)
Where is the multimodal goal post? On the Ability of Foundation Models to Recognize Contextually Important Moments
di: Surikuchi, Aditya K, et al.
Pubblicazione: (2026)
di: Surikuchi, Aditya K, et al.
Pubblicazione: (2026)
Hidden flaws behind expert-level accuracy of multimodal GPT-4 vision in medicine
di: Jin, Qiao, et al.
Pubblicazione: (2024)
di: Jin, Qiao, et al.
Pubblicazione: (2024)
Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models
di: Padlewski, Piotr, et al.
Pubblicazione: (2024)
di: Padlewski, Piotr, et al.
Pubblicazione: (2024)
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
di: Zhang, Ruixuan, et al.
Pubblicazione: (2025)
di: Zhang, Ruixuan, et al.
Pubblicazione: (2025)
What to align in multimodal contrastive learning?
di: Dufumier, Benoit, et al.
Pubblicazione: (2024)
di: Dufumier, Benoit, et al.
Pubblicazione: (2024)
Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence
di: Granite Vision Team, et al.
Pubblicazione: (2025)
di: Granite Vision Team, et al.
Pubblicazione: (2025)
Explaining latent representations of generative models with large multimodal models
di: Zhu, Mengdan, et al.
Pubblicazione: (2024)
di: Zhu, Mengdan, et al.
Pubblicazione: (2024)
LLM-grounded Video Diffusion Models
di: Lian, Long, et al.
Pubblicazione: (2023)
di: Lian, Long, et al.
Pubblicazione: (2023)
Revisiting Multi-Modal LLM Evaluation
di: Lu, Jian, et al.
Pubblicazione: (2024)
di: Lu, Jian, et al.
Pubblicazione: (2024)
FreeAct: Freeing Activations for LLM Quantization
di: Liu, Xiaohao, et al.
Pubblicazione: (2026)
di: Liu, Xiaohao, et al.
Pubblicazione: (2026)
Beyond Words: Multimodal LLM Knows When to Speak
di: Liao, Zikai, et al.
Pubblicazione: (2025)
di: Liao, Zikai, et al.
Pubblicazione: (2025)
VITA: Towards Open-Source Interactive Omni Multimodal LLM
di: Fu, Chaoyou, et al.
Pubblicazione: (2024)
di: Fu, Chaoyou, et al.
Pubblicazione: (2024)
TimeRefine: Temporal Grounding with Time Refining Video LLM
di: Wang, Xizi, et al.
Pubblicazione: (2024)
di: Wang, Xizi, et al.
Pubblicazione: (2024)
DIAMOND: An LLM-Driven Agent for Context-Aware Baseball Highlight Summarization
di: Kang, Jeonghun, et al.
Pubblicazione: (2025)
di: Kang, Jeonghun, et al.
Pubblicazione: (2025)
GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures
di: Patock, Jake R., et al.
Pubblicazione: (2025)
di: Patock, Jake R., et al.
Pubblicazione: (2025)
Qalam : A Multimodal LLM for Arabic Optical Character and Handwriting Recognition
di: Bhatia, Gagan, et al.
Pubblicazione: (2024)
di: Bhatia, Gagan, et al.
Pubblicazione: (2024)
OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
di: Ye, Hanrong, et al.
Pubblicazione: (2025)
di: Ye, Hanrong, et al.
Pubblicazione: (2025)
FaceLLM: A Multimodal Large Language Model for Face Understanding
di: Shahreza, Hatef Otroshi, et al.
Pubblicazione: (2025)
di: Shahreza, Hatef Otroshi, et al.
Pubblicazione: (2025)
PointLLM: Empowering Large Language Models to Understand Point Clouds
di: Xu, Runsen, et al.
Pubblicazione: (2023)
di: Xu, Runsen, et al.
Pubblicazione: (2023)
VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View
di: Schumann, Raphael, et al.
Pubblicazione: (2023)
di: Schumann, Raphael, et al.
Pubblicazione: (2023)
LLM-Optic: Unveiling the Capabilities of Large Language Models for Universal Visual Grounding
di: Zhao, Haoyu, et al.
Pubblicazione: (2024)
di: Zhao, Haoyu, et al.
Pubblicazione: (2024)
LLM-PCGC: Large Language Model-based Point Cloud Geometry Compression
di: Ye, Yuqi, et al.
Pubblicazione: (2024)
di: Ye, Yuqi, et al.
Pubblicazione: (2024)
VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions
di: Matsuda, Kazuki, et al.
Pubblicazione: (2025)
di: Matsuda, Kazuki, et al.
Pubblicazione: (2025)
Evaluating LLM -- Generated Multimodal Diagnosis from Medical Images and Symptom Analysis
di: Panagoulias, Dimitrios P., et al.
Pubblicazione: (2024)
di: Panagoulias, Dimitrios P., et al.
Pubblicazione: (2024)
From UAV Imagery to Agronomic Reasoning: A Multimodal LLM Benchmark for Plant Phenotyping
di: Wu, Yu, et al.
Pubblicazione: (2026)
di: Wu, Yu, et al.
Pubblicazione: (2026)
Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models
di: Horawalavithana, Sameera, et al.
Pubblicazione: (2026)
di: Horawalavithana, Sameera, et al.
Pubblicazione: (2026)
CRIMSON: A Clinically-Grounded LLM-Based Metric for Generative Radiology Report Evaluation
di: Baharoon, Mohammed, et al.
Pubblicazione: (2026)
di: Baharoon, Mohammed, et al.
Pubblicazione: (2026)
VISTA: Visual Integrated System for Tailored Automation in Math Problem Generation Using LLM
di: Lee, Jeongwoo, et al.
Pubblicazione: (2024)
di: Lee, Jeongwoo, et al.
Pubblicazione: (2024)
DriVLMe: Enhancing LLM-based Autonomous Driving Agents with Embodied and Social Experiences
di: Huang, Yidong, et al.
Pubblicazione: (2024)
di: Huang, Yidong, et al.
Pubblicazione: (2024)
VolDoGer: LLM-assisted Datasets for Domain Generalization in Vision-Language Tasks
di: Choi, Juhwan, et al.
Pubblicazione: (2024)
di: Choi, Juhwan, et al.
Pubblicazione: (2024)
A Survey of LLM-based Agents in Medicine: How far are we from Baymax?
di: Wang, Wenxuan, et al.
Pubblicazione: (2025)
di: Wang, Wenxuan, et al.
Pubblicazione: (2025)
Causal-SAM-LLM: Large Language Models as Causal Reasoners for Robust Medical Segmentation
di: Tang, Tao, et al.
Pubblicazione: (2025)
di: Tang, Tao, et al.
Pubblicazione: (2025)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
di: Wang, Ziyang, et al.
Pubblicazione: (2024)
di: Wang, Ziyang, et al.
Pubblicazione: (2024)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
di: Chen, Dongping, et al.
Pubblicazione: (2024)
di: Chen, Dongping, et al.
Pubblicazione: (2024)
Cost-effective Instruction Learning for Pathology Vision and Language Analysis
di: Chen, Kaitao, et al.
Pubblicazione: (2024)
di: Chen, Kaitao, et al.
Pubblicazione: (2024)
NavRAG: Generating User Demand Instructions for Embodied Navigation through Retrieval-Augmented LLM
di: Wang, Zihan, et al.
Pubblicazione: (2025)
di: Wang, Zihan, et al.
Pubblicazione: (2025)
StreetviewLLM: Extracting Geographic Information Using a Chain-of-Thought Multimodal Large Language Model
di: Li, Zongrong, et al.
Pubblicazione: (2024)
di: Li, Zongrong, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Automatic benchmarking of large multimodal models via iterative experiment programming
di: Conti, Alessandro, et al.
Pubblicazione: (2024) -
GlitchBench: Can large multimodal models detect video game glitches?
di: Taesiri, Mohammad Reza, et al.
Pubblicazione: (2023) -
MAIRA-1: A specialised large multimodal model for radiology report generation
di: Hyland, Stephanie L., et al.
Pubblicazione: (2023) -
Where is the multimodal goal post? On the Ability of Foundation Models to Recognize Contextually Important Moments
di: Surikuchi, Aditya K, et al.
Pubblicazione: (2026) -
Hidden flaws behind expert-level accuracy of multimodal GPT-4 vision in medicine
di: Jin, Qiao, et al.
Pubblicazione: (2024)