MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cui, Junbo, Xu, Bokai, Wang, Chongyi, Yu, Tianyu, Sun, Weiyue, Xu, Yingjing, Wang, Tianran, He, Zhihui, Ma, Wenshuo, Cai, Tianchi, Gui, Jiancheng, Zhang, Luoyuan, Sun, Xian, Huang, Fuwei, Chen, Moye, Lin, Zhuo, Liu, Hanyu, Gui, Qingxin, Han, Qingzhe, Wen, Yuyang, Liu, Huiping, Wang, Rongkang, Zhang, Yaqi, Wei, Hongliang, Chen, Chi, Li, You, Fang, Kechen, Zhou, Jie, Li, Yuxuan, Zeng, Guoyang, Xiao, Chaojun, Lin, Yankai, Han, Xu, Sun, Maosong, Liu, Zhiyuan, Yao, Yuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917450278240256
author Cui, Junbo
Xu, Bokai
Wang, Chongyi
Yu, Tianyu
Sun, Weiyue
Xu, Yingjing
Wang, Tianran
He, Zhihui
Ma, Wenshuo
Cai, Tianchi
Gui, Jiancheng
Zhang, Luoyuan
Sun, Xian
Huang, Fuwei
Chen, Moye
Lin, Zhuo
Liu, Hanyu
Gui, Qingxin
Han, Qingzhe
Wen, Yuyang
Liu, Huiping
Wang, Rongkang
Zhang, Yaqi
Wei, Hongliang
Chen, Chi
Li, You
Fang, Kechen
Zhou, Jie
Li, Yuxuan
Zeng, Guoyang
Xiao, Chaojun
Lin, Yankai
Han, Xu
Sun, Maosong
Liu, Zhiyuan
Yao, Yuan
author_facet Cui, Junbo
Xu, Bokai
Wang, Chongyi
Yu, Tianyu
Sun, Weiyue
Xu, Yingjing
Wang, Tianran
He, Zhihui
Ma, Wenshuo
Cai, Tianchi
Gui, Jiancheng
Zhang, Luoyuan
Sun, Xian
Huang, Fuwei
Chen, Moye
Lin, Zhuo
Liu, Hanyu
Gui, Qingxin
Han, Qingzhe
Wen, Yuyang
Liu, Huiping
Wang, Rongkang
Zhang, Yaqi
Wei, Hongliang
Chen, Chi
Li, You
Fang, Kechen
Zhou, Jie
Li, Yuxuan
Zeng, Guoyang
Xiao, Chaojun
Lin, Yankai
Han, Xu
Sun, Maosong
Liu, Zhiyuan
Yao, Yuan
contents Recent progress in multimodal large language models (MLLMs) has brought AI capabilities from static offline data processing to real-time streaming interaction, yet they still remain far from human-level multimodal interaction. The key bottlenecks are no longer modality coverage or latency alone, but the interaction paradigm itself. First, perception and response are still separated into alternating phases, preventing models from incorporating new inputs for timely adjustment during generation. Second, most current models remain reactive, responding only to explicit user requests instead of acting proactively in the evolving multimodal environment. We present MiniCPM-o 4.5, our latest effort towards human-like multimodal interaction, which mitigates these gaps by real-time full-duplex omni-modal interaction. It can see, listen, and speak simultaneously in real-time, while also exhibiting proactive behaviors such as issuing reminders or comments based on its continuous understanding of the live scene. The key technique behind MiniCPM-o 4.5 is Omni-Flow, a unified streaming framework that aligns omni-modal inputs and outputs along a shared temporal axis. This formulation converts conventional turn-based interaction into a full-duplex, time-aligned process, enabling simultaneous perception and response and allowing proactive behavior to arise within the same framework. With a total of 9B parameters, MiniCPM-o 4.5 approaches Gemini 2.5 Flash in vision-language capabilities, delivering state-of-the-art open-source performance at its scale. It also surpasses Qwen3-Omni-30B-A3B in omni-modal understanding and delivers better speech generation, with significantly higher computation efficiency. Driven by its efficient architecture design and inference optimization, the model can perform real-time full-duplex omni-modal interaction on edge devices with less than 12GB RAM cost.
format Preprint
id arxiv_https___arxiv_org_abs_2604_27393
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction
Cui, Junbo
Xu, Bokai
Wang, Chongyi
Yu, Tianyu
Sun, Weiyue
Xu, Yingjing
Wang, Tianran
He, Zhihui
Ma, Wenshuo
Cai, Tianchi
Gui, Jiancheng
Zhang, Luoyuan
Sun, Xian
Huang, Fuwei
Chen, Moye
Lin, Zhuo
Liu, Hanyu
Gui, Qingxin
Han, Qingzhe
Wen, Yuyang
Liu, Huiping
Wang, Rongkang
Zhang, Yaqi
Wei, Hongliang
Chen, Chi
Li, You
Fang, Kechen
Zhou, Jie
Li, Yuxuan
Zeng, Guoyang
Xiao, Chaojun
Lin, Yankai
Han, Xu
Sun, Maosong
Liu, Zhiyuan
Yao, Yuan
Computation and Language
Recent progress in multimodal large language models (MLLMs) has brought AI capabilities from static offline data processing to real-time streaming interaction, yet they still remain far from human-level multimodal interaction. The key bottlenecks are no longer modality coverage or latency alone, but the interaction paradigm itself. First, perception and response are still separated into alternating phases, preventing models from incorporating new inputs for timely adjustment during generation. Second, most current models remain reactive, responding only to explicit user requests instead of acting proactively in the evolving multimodal environment. We present MiniCPM-o 4.5, our latest effort towards human-like multimodal interaction, which mitigates these gaps by real-time full-duplex omni-modal interaction. It can see, listen, and speak simultaneously in real-time, while also exhibiting proactive behaviors such as issuing reminders or comments based on its continuous understanding of the live scene. The key technique behind MiniCPM-o 4.5 is Omni-Flow, a unified streaming framework that aligns omni-modal inputs and outputs along a shared temporal axis. This formulation converts conventional turn-based interaction into a full-duplex, time-aligned process, enabling simultaneous perception and response and allowing proactive behavior to arise within the same framework. With a total of 9B parameters, MiniCPM-o 4.5 approaches Gemini 2.5 Flash in vision-language capabilities, delivering state-of-the-art open-source performance at its scale. It also surpasses Qwen3-Omni-30B-A3B in omni-modal understanding and delivers better speech generation, with significantly higher computation efficiency. Driven by its efficient architecture design and inference optimization, the model can perform real-time full-duplex omni-modal interaction on edge devices with less than 12GB RAM cost.
title MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction
topic Computation and Language
url https://arxiv.org/abs/2604.27393