InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Dongchen, Sun, Yuyao, Zhang, Zilu, Huang, Leping, Zeng, Jianliang, Shu, Mao, Cao, Huo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
by: Chen, Zhe, et al.
Published: (2023)
by: Chen, Zhe, et al.
Published: (2023)
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
by: Wang, Weiyun, et al.
Published: (2025)
by: Wang, Weiyun, et al.
Published: (2025)
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
by: Zhu, Jinguo, et al.
Published: (2025)
by: Zhu, Jinguo, et al.
Published: (2025)
InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing
by: Tian, Changyao, et al.
Published: (2026)
by: Tian, Changyao, et al.
Published: (2026)
Driving with InternVL: Oustanding Champion in the Track on Driving with Language of the Autonomous Grand Challenge at CVPR 2024
by: Li, Jiahan, et al.
Published: (2024)
by: Li, Jiahan, et al.
Published: (2024)
Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
by: Gao, Zhangwei, et al.
Published: (2024)
by: Gao, Zhangwei, et al.
Published: (2024)
Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
by: Luo, Gen, et al.
Published: (2024)
by: Luo, Gen, et al.
Published: (2024)
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
by: Luo, Gen, et al.
Published: (2025)
by: Luo, Gen, et al.
Published: (2025)
FCoT-VL:Advancing Text-oriented Large Vision-Language Models with Efficient Visual Token Compression
by: Li, Jianjian, et al.
Published: (2025)
by: Li, Jianjian, et al.
Published: (2025)
VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language Models
by: Liang, Jiawei, et al.
Published: (2024)
by: Liang, Jiawei, et al.
Published: (2024)
DeepLatent: Think with Images via Parallel Latent Visual Reasoning
by: Lu, Dongchen, et al.
Published: (2026)
by: Lu, Dongchen, et al.
Published: (2026)
TinyChemVL: Advancing Chemical Vision-Language Models via Efficient Visual Token Reduction and Complex Reaction Tasks
by: Zhao, Xuanle, et al.
Published: (2025)
by: Zhao, Xuanle, et al.
Published: (2025)
InternVQA: Advancing Compressed Video Quality Assessment with Distilling Large Foundation Model
by: Guan, Fengbin, et al.
Published: (2025)
by: Guan, Fengbin, et al.
Published: (2025)
Singpath-VL Technical Report
by: Qiu, Zhen, et al.
Published: (2026)
by: Qiu, Zhen, et al.
Published: (2026)
Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs
by: Shu, Yan, et al.
Published: (2025)
by: Shu, Yan, et al.
Published: (2025)
ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification
by: He, Yefei, et al.
Published: (2024)
by: He, Yefei, et al.
Published: (2024)
Kimi-VL Technical Report
by: Kimi Team, et al.
Published: (2025)
by: Kimi Team, et al.
Published: (2025)
Kwai Keye-VL Technical Report
by: Kwai Keye Team, et al.
Published: (2025)
by: Kwai Keye Team, et al.
Published: (2025)
SAIL-VL2 Technical Report
by: Yin, Weijie, et al.
Published: (2025)
by: Yin, Weijie, et al.
Published: (2025)
VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration
by: Tu, Dezhan, et al.
Published: (2024)
by: Tu, Dezhan, et al.
Published: (2024)
Qwen3-VL Technical Report
by: Bai, Shuai, et al.
Published: (2025)
by: Bai, Shuai, et al.
Published: (2025)
Intern-S1: A Scientific Multimodal Foundation Model
by: Bai, Lei, et al.
Published: (2025)
by: Bai, Lei, et al.
Published: (2025)
Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
by: Wei, Zhixiang, et al.
Published: (2026)
by: Wei, Zhixiang, et al.
Published: (2026)
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Kwai Keye-VL 1.5 Technical Report
by: Yang, Biao, et al.
Published: (2025)
by: Yang, Biao, et al.
Published: (2025)
S1-VL: Scientific Multimodal Reasoning Model with Thinking-with-Images
by: Li, Qingxiao, et al.
Published: (2026)
by: Li, Qingxiao, et al.
Published: (2026)
Seed1.5-VL Technical Report
by: Guo, Dong, et al.
Published: (2025)
by: Guo, Dong, et al.
Published: (2025)
LiteMedCoT-VL: Parameter-Efficient Adaptation for Medical Visual Question Answering
by: Ma, Runze, et al.
Published: (2026)
by: Ma, Runze, et al.
Published: (2026)
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
by: Wang, Yi, et al.
Published: (2024)
by: Wang, Yi, et al.
Published: (2024)
InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
by: Wang, Haomin, et al.
Published: (2025)
by: Wang, Haomin, et al.
Published: (2025)
LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
by: Huang, Jiangyong, et al.
Published: (2025)
by: Huang, Jiangyong, et al.
Published: (2025)
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
STEP3-VL-10B Technical Report
by: Huang, Ailin, et al.
Published: (2026)
by: Huang, Ailin, et al.
Published: (2026)
HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices
by: HyperAI Team, et al.
Published: (2025)
by: HyperAI Team, et al.
Published: (2025)
Awaker2.5-VL: Stably Scaling MLLMs with Parameter-Efficient Mixture of Experts
by: Long, Jinqiang, et al.
Published: (2024)
by: Long, Jinqiang, et al.
Published: (2024)
ADEM-VL: Adaptive and Embedded Fusion for Efficient Vision-Language Tuning
by: Hao, Zhiwei, et al.
Published: (2024)
by: Hao, Zhiwei, et al.
Published: (2024)
Firebolt-VL: Efficient Vision-Language Understanding with Cross-Modality Modulation
by: Trinh, Quoc-Huy, et al.
Published: (2026)
by: Trinh, Quoc-Huy, et al.
Published: (2026)
VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
by: Yuan, Ruifeng, et al.
Published: (2025)
by: Yuan, Ruifeng, et al.
Published: (2025)
MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing
by: Wang, Wenjie, et al.
Published: (2026)
by: Wang, Wenjie, et al.
Published: (2026)
Intern-GS: Vision Model Guided Sparse-View 3D Gaussian Splatting
by: Sun, Xiangyu, et al.
Published: (2025)
by: Sun, Xiangyu, et al.
Published: (2025)
Similar Items
-
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
by: Chen, Zhe, et al.
Published: (2023) -
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
by: Wang, Weiyun, et al.
Published: (2025) -
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
by: Zhu, Jinguo, et al.
Published: (2025) -
InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing
by: Tian, Changyao, et al.
Published: (2026) -
Driving with InternVL: Oustanding Champion in the Track on Driving with Language of the Autonomous Grand Challenge at CVPR 2024
by: Li, Jiahan, et al.
Published: (2024)