Chat with AI: The Surprising Turn of Real-time Video Communication from Human to AI

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Jiangkai, Ren, Zhiyuan, Liu, Liming, Zhang, Xinggong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918215667417088
author Wu, Jiangkai
Ren, Zhiyuan
Liu, Liming
Zhang, Xinggong
author_facet Wu, Jiangkai
Ren, Zhiyuan
Liu, Liming
Zhang, Xinggong
contents AI Video Chat emerges as a new paradigm for Real-time Communication (RTC), where one peer is not a human, but a Multimodal Large Language Model (MLLM). This makes interaction between humans and AI more intuitive, as if chatting face-to-face with a real person. However, this poses significant challenges to latency, because the MLLM inference takes up most of the response time, leaving very little time for video streaming. Due to network uncertainty, transmission latency becomes a critical bottleneck preventing AI from being like a real person. To address this, we call for AI-oriented RTC research, exploring the network requirement shift from "humans watching video" to "AI understanding video". We begin by recognizing the main differences between AI Video Chat and traditional RTC. Then, through prototype measurements, we identify that ultra-low bitrate is a key factor for low latency. To reduce bitrate dramatically while maintaining MLLM accuracy, we propose Context-Aware Video Streaming that recognizes the importance of each video region for chat and allocates bitrate almost exclusively to chat-important regions. To evaluate the impact of video streaming quality on MLLM accuracy, we build the first benchmark, named Degraded Video Understanding Benchmark (DeViBench). Finally, we discuss some open questions and ongoing solutions for AI Video Chat. DeViBench is open-sourced at: https://github.com/pku-netvideo/DeViBench.
format Preprint
id arxiv_https___arxiv_org_abs_2507_10510
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Chat with AI: The Surprising Turn of Real-time Video Communication from Human to AI
Wu, Jiangkai
Ren, Zhiyuan
Liu, Liming
Zhang, Xinggong
Networking and Internet Architecture
Artificial Intelligence
Human-Computer Interaction
Multimedia
AI Video Chat emerges as a new paradigm for Real-time Communication (RTC), where one peer is not a human, but a Multimodal Large Language Model (MLLM). This makes interaction between humans and AI more intuitive, as if chatting face-to-face with a real person. However, this poses significant challenges to latency, because the MLLM inference takes up most of the response time, leaving very little time for video streaming. Due to network uncertainty, transmission latency becomes a critical bottleneck preventing AI from being like a real person. To address this, we call for AI-oriented RTC research, exploring the network requirement shift from "humans watching video" to "AI understanding video". We begin by recognizing the main differences between AI Video Chat and traditional RTC. Then, through prototype measurements, we identify that ultra-low bitrate is a key factor for low latency. To reduce bitrate dramatically while maintaining MLLM accuracy, we propose Context-Aware Video Streaming that recognizes the importance of each video region for chat and allocates bitrate almost exclusively to chat-important regions. To evaluate the impact of video streaming quality on MLLM accuracy, we build the first benchmark, named Degraded Video Understanding Benchmark (DeViBench). Finally, we discuss some open questions and ongoing solutions for AI Video Chat. DeViBench is open-sourced at: https://github.com/pku-netvideo/DeViBench.
title Chat with AI: The Surprising Turn of Real-time Video Communication from Human to AI
topic Networking and Internet Architecture
Artificial Intelligence
Human-Computer Interaction
Multimedia
url https://arxiv.org/abs/2507.10510