Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luo, Gen, Yang, Xue, Dou, Wenhan, Wang, Zhaokai, Liu, Jiawen, Dai, Jifeng, Qiao, Yu, Zhu, Xizhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
von: Luo, Gen, et al.
Veröffentlicht: (2025)
von: Luo, Gen, et al.
Veröffentlicht: (2025)
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
von: Chen, Zhe, et al.
Veröffentlicht: (2023)
von: Chen, Zhe, et al.
Veröffentlicht: (2023)
InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression
von: Lu, Dongchen, et al.
Veröffentlicht: (2025)
von: Lu, Dongchen, et al.
Veröffentlicht: (2025)
Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
von: Gao, Zhangwei, et al.
Veröffentlicht: (2024)
von: Gao, Zhangwei, et al.
Veröffentlicht: (2024)
InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing
von: Tian, Changyao, et al.
Veröffentlicht: (2026)
von: Tian, Changyao, et al.
Veröffentlicht: (2026)
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
von: Wang, Weiyun, et al.
Veröffentlicht: (2025)
von: Wang, Weiyun, et al.
Veröffentlicht: (2025)
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
von: Zhu, Jinguo, et al.
Veröffentlicht: (2025)
von: Zhu, Jinguo, et al.
Veröffentlicht: (2025)
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
Parameter-Inverted Image Pyramid Networks
von: Zhu, Xizhou, et al.
Veröffentlicht: (2024)
von: Zhu, Xizhou, et al.
Veröffentlicht: (2024)
Driving with InternVL: Oustanding Champion in the Track on Driving with Language of the Autonomous Grand Challenge at CVPR 2024
von: Li, Jiahan, et al.
Veröffentlicht: (2024)
von: Li, Jiahan, et al.
Veröffentlicht: (2024)
Auto MC-Reward: Automated Dense Reward Design with Large Language Models for Minecraft
von: Li, Hao, et al.
Veröffentlicht: (2023)
von: Li, Hao, et al.
Veröffentlicht: (2023)
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
von: Tao, Chenxin, et al.
Veröffentlicht: (2024)
von: Tao, Chenxin, et al.
Veröffentlicht: (2024)
GenExam: A Multidisciplinary Text-to-Image Exam
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
von: Luo, Gen, et al.
Veröffentlicht: (2025)
von: Luo, Gen, et al.
Veröffentlicht: (2025)
V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding
von: Ge, Junqi, et al.
Veröffentlicht: (2024)
von: Ge, Junqi, et al.
Veröffentlicht: (2024)
Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learning
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving
von: Cui, Erfei, et al.
Veröffentlicht: (2023)
von: Cui, Erfei, et al.
Veröffentlicht: (2023)
VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
von: Wang, Weiyun, et al.
Veröffentlicht: (2025)
von: Wang, Weiyun, et al.
Veröffentlicht: (2025)
Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs
von: Shu, Yan, et al.
Veröffentlicht: (2025)
von: Shu, Yan, et al.
Veröffentlicht: (2025)
NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
von: Tian, Changyao, et al.
Veröffentlicht: (2025)
von: Tian, Changyao, et al.
Veröffentlicht: (2025)
VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
von: Wu, Jiannan, et al.
Veröffentlicht: (2024)
von: Wu, Jiannan, et al.
Veröffentlicht: (2024)
CoMemo: LVLMs Need Image Context with Image Memory
von: Liu, Shi, et al.
Veröffentlicht: (2025)
von: Liu, Shi, et al.
Veröffentlicht: (2025)
Learning 1D Causal Visual Representation with De-focus Attention Networks
von: Tao, Chenxin, et al.
Veröffentlicht: (2024)
von: Tao, Chenxin, et al.
Veröffentlicht: (2024)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
von: Wang, Yi, et al.
Veröffentlicht: (2023)
von: Wang, Yi, et al.
Veröffentlicht: (2023)
Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
von: Duan, Yuchen, et al.
Veröffentlicht: (2024)
von: Duan, Yuchen, et al.
Veröffentlicht: (2024)
Hierarchical Multimodal Pre-training for Visually Rich Webpage Understanding
von: Xu, Hongshen, et al.
Veröffentlicht: (2024)
von: Xu, Hongshen, et al.
Veröffentlicht: (2024)
InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
von: Wang, Haomin, et al.
Veröffentlicht: (2025)
von: Wang, Haomin, et al.
Veröffentlicht: (2025)
MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings
von: Chen, Haonan, et al.
Veröffentlicht: (2025)
von: Chen, Haonan, et al.
Veröffentlicht: (2025)
MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
von: MiroMind Team, et al.
Veröffentlicht: (2025)
von: MiroMind Team, et al.
Veröffentlicht: (2025)
SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence
von: Gong, Ziyang, et al.
Veröffentlicht: (2025)
von: Gong, Ziyang, et al.
Veröffentlicht: (2025)
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
Design as Desired: Utilizing Visual Question Answering for Multimodal Pre-training
von: Su, Tongkun, et al.
Veröffentlicht: (2024)
von: Su, Tongkun, et al.
Veröffentlicht: (2024)
InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models
von: Deng, Nianchen, et al.
Veröffentlicht: (2025)
von: Deng, Nianchen, et al.
Veröffentlicht: (2025)
Wan-Image: Pushing the Boundaries of Generative Visual Intelligence
von: Mao, Chaojie, et al.
Veröffentlicht: (2026)
von: Mao, Chaojie, et al.
Veröffentlicht: (2026)
VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language Models
von: Liang, Jiawei, et al.
Veröffentlicht: (2024)
von: Liang, Jiawei, et al.
Veröffentlicht: (2024)
Intern-S1: A Scientific Multimodal Foundation Model
von: Bai, Lei, et al.
Veröffentlicht: (2025)
von: Bai, Lei, et al.
Veröffentlicht: (2025)
Pushing Boundaries: Exploring Zero Shot Object Classification with Large Multimodal Models
von: Islam, Ashhadul, et al.
Veröffentlicht: (2023)
von: Islam, Ashhadul, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
von: Luo, Gen, et al.
Veröffentlicht: (2025) -
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
von: Chen, Zhe, et al.
Veröffentlicht: (2023) -
InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression
von: Lu, Dongchen, et al.
Veröffentlicht: (2025) -
Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
von: Gao, Zhangwei, et al.
Veröffentlicht: (2024) -
InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing
von: Tian, Changyao, et al.
Veröffentlicht: (2026)