CogVLM2: Visual Language Models for Image and Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Wenyi, Wang, Weihan, Ding, Ming, Yu, Wenmeng, Lv, Qingsong, Wang, Yan, Cheng, Yean, Huang, Shiyu, Ji, Junhui, Xue, Zhao, Zhao, Lei, Yang, Zhuoyi, Gu, Xiaotao, Zhang, Xiaohan, Feng, Guanyu, Yin, Da, Wang, Zihan, Qi, Ji, Song, Xixuan, Zhang, Peng, Liu, Debing, Xu, Bin, Li, Juanzi, Dong, Yuxiao, Tang, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CogVLM: Visual Expert for Pretrained Language Models
by: Wang, Weihan, et al.
Published: (2023)
by: Wang, Weihan, et al.
Published: (2023)
CogAgent: A Visual Language Model for GUI Agents
by: Hong, Wenyi, et al.
Published: (2023)
by: Hong, Wenyi, et al.
Published: (2023)
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
by: Yang, Zhuoyi, et al.
Published: (2024)
by: Yang, Zhuoyi, et al.
Published: (2024)
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
by: Hong, Wenyi, et al.
Published: (2025)
by: Hong, Wenyi, et al.
Published: (2025)
CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning
by: Qi, Ji, et al.
Published: (2024)
by: Qi, Ji, et al.
Published: (2024)
LVBench: An Extreme Long Video Understanding Benchmark
by: Wang, Weihan, et al.
Published: (2024)
by: Wang, Weihan, et al.
Published: (2024)
CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion
by: Zheng, Wendi, et al.
Published: (2024)
by: Zheng, Wendi, et al.
Published: (2024)
HPE-CogVLM: Advancing Vision Language Models with a Head Pose Grounding Task
by: Tian, Yu, et al.
Published: (2024)
by: Tian, Yu, et al.
Published: (2024)
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
by: Wu, Yuhang, et al.
Published: (2024)
by: Wu, Yuhang, et al.
Published: (2024)
GLM-OCR Technical Report
by: Duan, Shuaiqi, et al.
Published: (2026)
by: Duan, Shuaiqi, et al.
Published: (2026)
MulCogBench: A Multi-modal Cognitive Benchmark Dataset for Evaluating Chinese and English Computational Language Models
by: Zhang, Yunhao, et al.
Published: (2024)
by: Zhang, Yunhao, et al.
Published: (2024)
MathGLM-Vision: Solving Mathematical Problems with Multi-Modal Large Language Model
by: Yang, Zhen, et al.
Published: (2024)
by: Yang, Zhen, et al.
Published: (2024)
The Effect of Gamification on Employee Boredom and Performance*
by: Zhuoyi Zhao
Published: (2024)
by: Zhuoyi Zhao
Published: (2024)
UI2Code^N: UI-to-Code Generation as Interactive Visual Optimization
by: Yang, Zhen, et al.
Published: (2025)
by: Yang, Zhen, et al.
Published: (2025)
TraceRAG: A LLM-Based Framework for Explainable Android Malware Detection and Behavior Analysis
by: Zhang, Guangyu, et al.
Published: (2025)
by: Zhang, Guangyu, et al.
Published: (2025)
Extensive Self-Contrast Enables Feedback-Free Language Model Alignment
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
Rethinking the Text-Vision Reasoning Imbalance in MLLMs through the Lens of Training Recipes
by: Yao, Guanyu, et al.
Published: (2025)
by: Yao, Guanyu, et al.
Published: (2025)
NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts
by: Zhang, Shudan, et al.
Published: (2024)
by: Zhang, Shudan, et al.
Published: (2024)
LongAlign: A Recipe for Long Context Alignment of Large Language Models
by: Bai, Yushi, et al.
Published: (2024)
by: Bai, Yushi, et al.
Published: (2024)
ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline
by: Xu, Yifan, et al.
Published: (2024)
by: Xu, Yifan, et al.
Published: (2024)
Electric-Field-Tunable Luttinger compensated antiferromagnetism in double CrCl2 chains
by: Guo, Deping, et al.
Published: (2025)
by: Guo, Deping, et al.
Published: (2025)
Identifiability and Consistent Estimation for Gaussian Chain Graph Models
by: Zhao, Ruixuan, et al.
Published: (2023)
by: Zhao, Ruixuan, et al.
Published: (2023)
Untie the Knots: An Efficient Data Augmentation Strategy for Long-Context Pre-Training in Language Models
by: Tian, Junfeng, et al.
Published: (2024)
by: Tian, Junfeng, et al.
Published: (2024)
NITP: Next Implicit Token Prediction for LLM Pre-training
by: Zhang, Xiangdong, et al.
Published: (2026)
by: Zhang, Xiangdong, et al.
Published: (2026)
Multi-Task Multi-Agent Reinforcement Learning via Skill Graphs
by: Zhu, Guobin, et al.
Published: (2025)
by: Zhu, Guobin, et al.
Published: (2025)
Investigation into the overcurrent failure and combustion characteristics of copper‐clad aluminum conductors
by: Weifeng Wang, et al.
Published: (2024)
by: Weifeng Wang, et al.
Published: (2024)
CogDrive: Cognition-Driven Multimodal Prediction-Planning Fusion for Safe Autonomy
by: Huang, Heye, et al.
Published: (2025)
by: Huang, Heye, et al.
Published: (2025)
CogPlanner: Unveiling the Potential of Agentic Multimodal Retrieval Augmented Generation with Planning
by: Yu, Xiaohan, et al.
Published: (2025)
by: Yu, Xiaohan, et al.
Published: (2025)
DreamPolish: Domain Score Distillation With Progressive Geometry Generation
by: Cheng, Yean, et al.
Published: (2024)
by: Cheng, Yean, et al.
Published: (2024)
Toward Practical Age-of-Information Scheduling in 5G Cellular
by: Zhao, Zhuoyi, et al.
Published: (2026)
by: Zhao, Zhuoyi, et al.
Published: (2026)
Quantum symmetric pairs via Hall algebras
by: Lu, Ming, et al.
Published: (2025)
by: Lu, Ming, et al.
Published: (2025)
Optimizing Age of Information without Knowing the Age of Information
by: Zhao, Zhuoyi, et al.
Published: (2025)
by: Zhao, Zhuoyi, et al.
Published: (2025)
Erianin Protects Human Umbilical Vein Endothelial Cells From Oxidized Low‐Density Lipoprotein‐Induced Apoptosis and Oxidative Stress Through Activation of Nuclear Factor E2‐Related Factor 2 Signaling
by: Zhaowei Wang, et al.
Published: (2025)
by: Zhaowei Wang, et al.
Published: (2025)
Regularization Prescription for the Mixing Between Nonlocal Gluon and Quark Operators
by: Ji, Yao, et al.
Published: (2025)
by: Ji, Yao, et al.
Published: (2025)
Tunable altermagnetism via inter-chain engineering in parallelassembled atomic chains
by: Guo, Deping, et al.
Published: (2025)
by: Guo, Deping, et al.
Published: (2025)
MLLM-Selector: Necessity and Diversity-driven High-Value Data Selection for Enhanced Visual Instruction Tuning
by: Ma, Yiwei, et al.
Published: (2025)
by: Ma, Yiwei, et al.
Published: (2025)
CogStream: Context-guided Streaming Video Question Answering
by: Zhao, Zicheng, et al.
Published: (2025)
by: Zhao, Zicheng, et al.
Published: (2025)
UrbanVLP: Multi-Granularity Vision-Language Pretraining for Urban Socioeconomic Indicator Prediction
by: Hao, Xixuan, et al.
Published: (2024)
by: Hao, Xixuan, et al.
Published: (2024)
SedarEval: Automated Evaluation using Self-Adaptive Rubrics
by: Fan, Zhiyuan, et al.
Published: (2025)
by: Fan, Zhiyuan, et al.
Published: (2025)
Inf-DiT: Upsampling Any-Resolution Image with Memory-Efficient Diffusion Transformer
by: Yang, Zhuoyi, et al.
Published: (2024)
by: Yang, Zhuoyi, et al.
Published: (2024)
Similar Items
-
CogVLM: Visual Expert for Pretrained Language Models
by: Wang, Weihan, et al.
Published: (2023) -
CogAgent: A Visual Language Model for GUI Agents
by: Hong, Wenyi, et al.
Published: (2023) -
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
by: Yang, Zhuoyi, et al.
Published: (2024) -
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
by: Hong, Wenyi, et al.
Published: (2025) -
CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning
by: Qi, Ji, et al.
Published: (2024)