CogVLM2: Visual Language Models for Image and Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hong, Wenyi, Wang, Weihan, Ding, Ming, Yu, Wenmeng, Lv, Qingsong, Wang, Yan, Cheng, Yean, Huang, Shiyu, Ji, Junhui, Xue, Zhao, Zhao, Lei, Yang, Zhuoyi, Gu, Xiaotao, Zhang, Xiaohan, Feng, Guanyu, Yin, Da, Wang, Zihan, Qi, Ji, Song, Xixuan, Zhang, Peng, Liu, Debing, Xu, Bin, Li, Juanzi, Dong, Yuxiao, Tang, Jie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913485514866688
author Hong, Wenyi
Wang, Weihan
Ding, Ming
Yu, Wenmeng
Lv, Qingsong
Wang, Yan
Cheng, Yean
Huang, Shiyu
Ji, Junhui
Xue, Zhao
Zhao, Lei
Yang, Zhuoyi
Gu, Xiaotao
Zhang, Xiaohan
Feng, Guanyu
Yin, Da
Wang, Zihan
Qi, Ji
Song, Xixuan
Zhang, Peng
Liu, Debing
Xu, Bin
Li, Juanzi
Dong, Yuxiao
Tang, Jie
author_facet Hong, Wenyi
Wang, Weihan
Ding, Ming
Yu, Wenmeng
Lv, Qingsong
Wang, Yan
Cheng, Yean
Huang, Shiyu
Ji, Junhui
Xue, Zhao
Zhao, Lei
Yang, Zhuoyi
Gu, Xiaotao
Zhang, Xiaohan
Feng, Guanyu
Yin, Da
Wang, Zihan
Qi, Ji
Song, Xixuan
Zhang, Peng
Liu, Debing
Xu, Bin
Li, Juanzi
Dong, Yuxiao
Tang, Jie
contents Beginning with VisualGLM and CogVLM, we are continuously exploring VLMs in pursuit of enhanced vision-language fusion, efficient higher-resolution architecture, and broader modalities and applications. Here we propose the CogVLM2 family, a new generation of visual language models for image and video understanding including CogVLM2, CogVLM2-Video and GLM-4V. As an image understanding model, CogVLM2 inherits the visual expert architecture with improved training recipes in both pre-training and post-training stages, supporting input resolution up to $1344 \times 1344$ pixels. As a video understanding model, CogVLM2-Video integrates multi-frame input with timestamps and proposes automated temporal grounding data construction. Notably, CogVLM2 family has achieved state-of-the-art results on benchmarks like MMBench, MM-Vet, TextVQA, MVBench and VCGBench. All models are open-sourced in https://github.com/THUDM/CogVLM2 and https://github.com/THUDM/GLM-4, contributing to the advancement of the field.
format Preprint
id arxiv_https___arxiv_org_abs_2408_16500
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CogVLM2: Visual Language Models for Image and Video Understanding
Hong, Wenyi
Wang, Weihan
Ding, Ming
Yu, Wenmeng
Lv, Qingsong
Wang, Yan
Cheng, Yean
Huang, Shiyu
Ji, Junhui
Xue, Zhao
Zhao, Lei
Yang, Zhuoyi
Gu, Xiaotao
Zhang, Xiaohan
Feng, Guanyu
Yin, Da
Wang, Zihan
Qi, Ji
Song, Xixuan
Zhang, Peng
Liu, Debing
Xu, Bin
Li, Juanzi
Dong, Yuxiao
Tang, Jie
Computer Vision and Pattern Recognition
Beginning with VisualGLM and CogVLM, we are continuously exploring VLMs in pursuit of enhanced vision-language fusion, efficient higher-resolution architecture, and broader modalities and applications. Here we propose the CogVLM2 family, a new generation of visual language models for image and video understanding including CogVLM2, CogVLM2-Video and GLM-4V. As an image understanding model, CogVLM2 inherits the visual expert architecture with improved training recipes in both pre-training and post-training stages, supporting input resolution up to $1344 \times 1344$ pixels. As a video understanding model, CogVLM2-Video integrates multi-frame input with timestamps and proposes automated temporal grounding data construction. Notably, CogVLM2 family has achieved state-of-the-art results on benchmarks like MMBench, MM-Vet, TextVQA, MVBench and VCGBench. All models are open-sourced in https://github.com/THUDM/CogVLM2 and https://github.com/THUDM/GLM-4, contributing to the advancement of the field.
title CogVLM2: Visual Language Models for Image and Video Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.16500