PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Xudong, Guan, Huankang, Bo, Yang, Chen, Jinpeng, Guo, Xintong, Li, Shuhan, Liu, Fang, Sun, Peiwen, Li, Xueying, Zhang, Wei, Yang, Xue, Liu, Rui, Li, Hongsheng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914293517123584
author Lu, Xudong
Guan, Huankang
Bo, Yang
Chen, Jinpeng
Guo, Xintong
Li, Shuhan
Liu, Fang
Sun, Peiwen
Li, Xueying
Zhang, Wei
Yang, Xue
Liu, Rui
Li, Hongsheng
author_facet Lu, Xudong
Guan, Huankang
Bo, Yang
Chen, Jinpeng
Guo, Xintong
Li, Shuhan
Liu, Fang
Sun, Peiwen
Li, Xueying
Zhang, Wei
Yang, Xue
Liu, Rui
Li, Hongsheng
contents Multimodal Large Language Models excel at offline audio-visual understanding, but their ability to serve as mobile assistants in continuous real-world streams remains underexplored. In daily phone use, mobile assistants must track streaming audio-visual inputs and respond at the right time, yet existing benchmarks are often restricted to multiple-choice questions or use shorter videos. In this paper, we introduce PhoStream, the first mobile-centric streaming benchmark that unifies on-screen and off-screen scenarios to evaluate video, audio, and temporal reasoning. PhoStream contains 5,572 open-ended QA pairs from 578 videos across 4 scenarios and 10 capabilities. We build it with an Automated Generative Pipeline backed by rigorous human verification, and evaluate models using a realistic Online Inference Pipeline and LLM-as-a-Judge evaluation for open-ended responses. Experiments reveal a temporal asymmetry in LLM-judged scores (0-100): models perform well on Instant and Backward tasks (Gemini 3 Pro exceeds 80), but drop sharply on Forward tasks (16.40), largely due to early responses before the required visual and audio cues appear. This highlights a fundamental limitation: current MLLMs struggle to decide when to speak, not just what to say. Code and datasets used in this work will be made publicly accessible at https://github.com/Lucky-Lance/PhoStream.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22575
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios
Lu, Xudong
Guan, Huankang
Bo, Yang
Chen, Jinpeng
Guo, Xintong
Li, Shuhan
Liu, Fang
Sun, Peiwen
Li, Xueying
Zhang, Wei
Yang, Xue
Liu, Rui
Li, Hongsheng
Computer Vision and Pattern Recognition
Computation and Language
Multimodal Large Language Models excel at offline audio-visual understanding, but their ability to serve as mobile assistants in continuous real-world streams remains underexplored. In daily phone use, mobile assistants must track streaming audio-visual inputs and respond at the right time, yet existing benchmarks are often restricted to multiple-choice questions or use shorter videos. In this paper, we introduce PhoStream, the first mobile-centric streaming benchmark that unifies on-screen and off-screen scenarios to evaluate video, audio, and temporal reasoning. PhoStream contains 5,572 open-ended QA pairs from 578 videos across 4 scenarios and 10 capabilities. We build it with an Automated Generative Pipeline backed by rigorous human verification, and evaluate models using a realistic Online Inference Pipeline and LLM-as-a-Judge evaluation for open-ended responses. Experiments reveal a temporal asymmetry in LLM-judged scores (0-100): models perform well on Instant and Backward tasks (Gemini 3 Pro exceeds 80), but drop sharply on Forward tasks (16.40), largely due to early responses before the required visual and audio cues appear. This highlights a fundamental limitation: current MLLMs struggle to decide when to speak, not just what to say. Code and datasets used in this work will be made publicly accessible at https://github.com/Lucky-Lance/PhoStream.
title PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2601.22575