MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, Yuan, Yu, Tianyu, Zhang, Ao, Wang, Chongyi, Cui, Junbo, Zhu, Hongji, Cai, Tianchi, Li, Haoyu, Zhao, Weilin, He, Zhihui, Chen, Qianyu, Zhou, Huarong, Zou, Zhensheng, Zhang, Haoye, Hu, Shengding, Zheng, Zhi, Zhou, Jie, Cai, Jie, Han, Xu, Zeng, Guoyang, Li, Dahai, Liu, Zhiyuan, Sun, Maosong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
by: Hu, Shengding, et al.
Published: (2024)
by: Hu, Shengding, et al.
Published: (2024)
MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
by: Yu, Tianyu, et al.
Published: (2025)
by: Yu, Tianyu, et al.
Published: (2025)
MiniCPM4: Ultra-Efficient LLMs on End Devices
by: MiniCPM Team, et al.
Published: (2025)
by: MiniCPM Team, et al.
Published: (2025)
MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction
by: Cui, Junbo, et al.
Published: (2026)
by: Cui, Junbo, et al.
Published: (2026)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
by: MiniCPM Team, et al.
Published: (2026)
by: MiniCPM Team, et al.
Published: (2026)
Densing Law of LLMs
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
DecorateLM: Data Engineering through Corpus Rating, Tagging, and Editing with Language Models
by: Zhao, Ranchi, et al.
Published: (2024)
by: Zhao, Ranchi, et al.
Published: (2024)
AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning
by: Zhang, Zhong, et al.
Published: (2025)
by: Zhang, Zhong, et al.
Published: (2025)
Predicting Emergent Abilities with Infinite Resolution Evaluation
by: Hu, Shengding, et al.
Published: (2023)
by: Hu, Shengding, et al.
Published: (2023)
LLM$\times$MapReduce-V3: Enabling Interactive In-Depth Survey Generation through a MCP-Driven Hierarchically Modular Agent System
by: Chao, Yu, et al.
Published: (2025)
by: Chao, Yu, et al.
Published: (2025)
Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages
by: Hu, Jinyi, et al.
Published: (2023)
by: Hu, Jinyi, et al.
Published: (2023)
States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
by: Chen, Junhao, et al.
Published: (2024)
by: Chen, Junhao, et al.
Published: (2024)
UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
by: Gu, Tiancheng, et al.
Published: (2025)
by: Gu, Tiancheng, et al.
Published: (2025)
Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models
by: Zhang, Xinrong, et al.
Published: (2024)
by: Zhang, Xinrong, et al.
Published: (2024)
UltraEval: A Lightweight Platform for Flexible and Comprehensive Evaluation for LLMs
by: He, Chaoqun, et al.
Published: (2024)
by: He, Chaoqun, et al.
Published: (2024)
Experimental and numerical investigations on formation mechanisms of hollow structures in overflow water‐assisted injection‐molded parts of short‐glass‐fiber‐reinforced polypropylene
by: Qingsong Jiang, et al.
Published: (2024)
by: Qingsong Jiang, et al.
Published: (2024)
A hypergraph bipartite Turán problem with odd uniformity
by: Ma, Jie, et al.
Published: (2024)
by: Ma, Jie, et al.
Published: (2024)
Stuffed Mamba: Oversized States Lead to the Inability to Forget
by: Chen, Yingfa, et al.
Published: (2024)
by: Chen, Yingfa, et al.
Published: (2024)
VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
by: Zhou, Yixuan, et al.
Published: (2025)
by: Zhou, Yixuan, et al.
Published: (2025)
Unified View of Grokking, Double Descent and Emergent Abilities: A Perspective from Circuits Competition
by: Huang, Yufei, et al.
Published: (2024)
by: Huang, Yufei, et al.
Published: (2024)
NPSVC++: Nonparallel Classifiers Encounter Representation Learning
by: Zhang, Junhong, et al.
Published: (2024)
by: Zhang, Junhong, et al.
Published: (2024)
Percutaneous Transhepatic Cholangioscopy in Hepatolithiasis Associated With Decompensated Cirrhosis: A Retrospective Cohort Study
by: Qianyu Yan, et al.
Published: (2024)
by: Qianyu Yan, et al.
Published: (2024)
Naringenin Inhibits Ferroptosis in Renal Tubular Epithelial Cells of Diabetic Nephropathy Through SIRT1/FOXO3a Signaling Pathway
by: Yi Zhou, et al.
Published: (2025)
by: Yi Zhou, et al.
Published: (2025)
AgentCPM-Report: Interleaving Drafting and Deepening for Open-Ended Deep Research
by: Li, Yishan, et al.
Published: (2026)
by: Li, Yishan, et al.
Published: (2026)
MiniPLM: Knowledge Distillation for Pre-Training Language Models
by: Gu, Yuxian, et al.
Published: (2024)
by: Gu, Yuxian, et al.
Published: (2024)
DeTRAP: RISC-V Return Address Protection With Debug Triggers
by: Richter, Isaac, et al.
Published: (2024)
by: Richter, Isaac, et al.
Published: (2024)
TopV-Nav: Unlocking the Top-View Spatial Reasoning Potential of MLLM for Zero-shot Object Navigation
by: Zhong, Linqing, et al.
Published: (2024)
by: Zhong, Linqing, et al.
Published: (2024)
A Global-Local Graph Attention Network for Traffic Forecasting
by: Zhang, Tianchi
Published: (2026)
by: Zhang, Tianchi
Published: (2026)
RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness
by: Yu, Tianyu, et al.
Published: (2024)
by: Yu, Tianyu, et al.
Published: (2024)
Generic effective sources for first-order in mass-ratio gravitational self-force calculations in Schwarzschild spacetime
by: Zhang, Chao, et al.
Published: (2025)
by: Zhang, Chao, et al.
Published: (2025)
An introduction to $V$-filtrations
by: Chen, Qianyu, et al.
Published: (2024)
by: Chen, Qianyu, et al.
Published: (2024)
Assessing the Impact of Artificial Intelligence and Green Finance on Energy Efficiency: Based on Super‐Efficiency SBM and Tobit Two‐Stage Models
by: Hongji Zhou, et al.
Published: (2025)
by: Hongji Zhou, et al.
Published: (2025)
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
by: Huang, Yuxiang, et al.
Published: (2025)
by: Huang, Yuxiang, et al.
Published: (2025)
Mechanism of Strength–Toughness Balance in 2Cr13MoV Martensitic Stainless Steel under Short‐Time Heat Treatment
by: Yong‐yong Jia, et al.
Published: (2025)
by: Yong‐yong Jia, et al.
Published: (2025)
PhoneWorld: Scaling Phone-Use Agent Environments
by: Tang, Zhengyang, et al.
Published: (2026)
by: Tang, Zhengyang, et al.
Published: (2026)
CPM and PERT in Library Management.
by: Main, Linda
Published: (1989)
by: Main, Linda
Published: (1989)
Some exact results on $4$-cycles: stability and supersaturation
by: He, Jialin, et al.
Published: (2019)
by: He, Jialin, et al.
Published: (2019)
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
by: Yu, Tianyu, et al.
Published: (2023)
by: Yu, Tianyu, et al.
Published: (2023)
Efficiency-Aware Computational Intelligence for Resource-Constrained Manufacturing Toward Edge-Ready Deployment
by: Zhou, Qianyu
Published: (2025)
by: Zhou, Qianyu
Published: (2025)
AgentCPM-Explore: Realizing Long-Horizon Deep Exploration for Edge-Scale Agents
by: Chen, Haotian, et al.
Published: (2026)
by: Chen, Haotian, et al.
Published: (2026)
Similar Items
-
MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
by: Hu, Shengding, et al.
Published: (2024) -
MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
by: Yu, Tianyu, et al.
Published: (2025) -
MiniCPM4: Ultra-Efficient LLMs on End Devices
by: MiniCPM Team, et al.
Published: (2025) -
MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction
by: Cui, Junbo, et al.
Published: (2026) -
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
by: MiniCPM Team, et al.
Published: (2026)