IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jinzhao, Chen, Yinuo, Song, Wenxuan, Lei, Yijia, Zhang, Yichi, Yan, Honglei, Pan, Panwang, Liu, Miao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy
by: Li, Jinzhao, et al.
Published: (2026)
by: Li, Jinzhao, et al.
Published: (2026)
PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers
by: Lin, Yuchen, et al.
Published: (2025)
by: Lin, Yuchen, et al.
Published: (2025)
MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second
by: Lin, Chenguo, et al.
Published: (2025)
by: Lin, Chenguo, et al.
Published: (2025)
X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding
by: Sun, Peiwen, et al.
Published: (2026)
by: Sun, Peiwen, et al.
Published: (2026)
MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning
by: Wang, Yueqian, et al.
Published: (2025)
by: Wang, Yueqian, et al.
Published: (2025)
JarvisIR: Elevating Autonomous Driving Perception with Intelligent Image Restoration
by: Lin, Yunlong, et al.
Published: (2025)
by: Lin, Yunlong, et al.
Published: (2025)
StrLoRA: Towards Streaming Continual Visual Instruction Tuning for MLLMs
by: Che, Chang, et al.
Published: (2026)
by: Che, Chang, et al.
Published: (2026)
ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
by: Wang, Yueqian, et al.
Published: (2025)
by: Wang, Yueqian, et al.
Published: (2025)
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
by: Lin, Junming, et al.
Published: (2024)
by: Lin, Junming, et al.
Published: (2024)
Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models
by: Pan, Panwang, et al.
Published: (2025)
by: Pan, Panwang, et al.
Published: (2025)
CurveStream: Boosting Streaming Video Understanding in MLLMs via Curvature-Aware Hierarchical Visual Memory Management
by: Wang, Chao, et al.
Published: (2026)
by: Wang, Chao, et al.
Published: (2026)
EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams
by: Ran, Dongchuan, et al.
Published: (2026)
by: Ran, Dongchuan, et al.
Published: (2026)
LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs
by: Chen, Jingfeng, et al.
Published: (2026)
by: Chen, Jingfeng, et al.
Published: (2026)
JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent
by: Lin, Yunlong, et al.
Published: (2025)
by: Lin, Yunlong, et al.
Published: (2025)
Vision-Proprioception Fusion with Mamba2 in End-to-End Reinforcement Learning for Motion Control
by: Tao, Xiaowen, et al.
Published: (2025)
by: Tao, Xiaowen, et al.
Published: (2025)
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
Memory Helps, but Confabulation Misleads: Understanding Streaming Events in Videos with MLLMs
by: Zhang, Gengyuan, et al.
Published: (2025)
by: Zhang, Gengyuan, et al.
Published: (2025)
DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation
by: Lin, Chenguo, et al.
Published: (2025)
by: Lin, Chenguo, et al.
Published: (2025)
InstructLayout: Instruction-Driven 2D and 3D Layout Synthesis with Semantic Graph Prior
by: Lin, Chenguo, et al.
Published: (2024)
by: Lin, Chenguo, et al.
Published: (2024)
Reasoning over Video: Evaluating How MLLMs Extract, Integrate, and Reconstruct Spatiotemporal Evidence
by: Bang, Seunghwan, et al.
Published: (2026)
by: Bang, Seunghwan, et al.
Published: (2026)
OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams
by: Yan, Yibin, et al.
Published: (2026)
by: Yan, Yibin, et al.
Published: (2026)
Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video
by: Zhang, Yulin, et al.
Published: (2025)
by: Zhang, Yulin, et al.
Published: (2025)
Seeing the Unseen in Low-light Spike Streams
by: Hu, Liwen, et al.
Published: (2025)
by: Hu, Liwen, et al.
Published: (2025)
RO-Bench: Large-scale robustness evaluation of MLLMs with text-driven counterfactual videos
by: Yang, Zixi, et al.
Published: (2025)
by: Yang, Zixi, et al.
Published: (2025)
StreamPro: From Reactive Perception to Proactive Decision-Making in Streaming Video
by: Li, Ao, et al.
Published: (2026)
by: Li, Ao, et al.
Published: (2026)
Ground What You See: Hallucination-Resistant MLLMs via Caption Feedback, Diversity-Aware Sampling, and Conflict Regularization
by: Pan, Miao, et al.
Published: (2026)
by: Pan, Miao, et al.
Published: (2026)
MokA: Multimodal Low-Rank Adaptation for MLLMs
by: Wei, Yake, et al.
Published: (2025)
by: Wei, Yake, et al.
Published: (2025)
LOP: Learning Optimal Pruning for Efficient On-Demand MLLMs Scaling
by: Zhang, Zhihan, et al.
Published: (2025)
by: Zhang, Zhihan, et al.
Published: (2025)
On the Generalization Capacities of MLLMs for Spatial Intelligence
by: Zhang, Gongjie, et al.
Published: (2026)
by: Zhang, Gongjie, et al.
Published: (2026)
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
by: Zhao, Ruixiang, et al.
Published: (2026)
by: Zhao, Ruixiang, et al.
Published: (2026)
iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA
by: Zhao, Zhaoran, et al.
Published: (2025)
by: Zhao, Zhaoran, et al.
Published: (2025)
When MLLMs Meet Compression Distortion: A Coding Paradigm Tailored to MLLMs
by: Liu, Jinming, et al.
Published: (2025)
by: Liu, Jinming, et al.
Published: (2025)
Pose-RFT: Enhancing MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning
by: Li, Bao, et al.
Published: (2025)
by: Li, Bao, et al.
Published: (2025)
SuperMat: Physically Consistent PBR Material Estimation at Interactive Rates
by: Hong, Yijia, et al.
Published: (2024)
by: Hong, Yijia, et al.
Published: (2024)
Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
by: Tang, Yuqi, et al.
Published: (2026)
by: Tang, Yuqi, et al.
Published: (2026)
Proactive Scene Decomposition and Reconstruction
by: Li, Baicheng, et al.
Published: (2025)
by: Li, Baicheng, et al.
Published: (2025)
Depth-agnostic Single Image Dehazing
by: Xu, Honglei, et al.
Published: (2024)
by: Xu, Honglei, et al.
Published: (2024)
InterGen: Diffusion-based Multi-human Motion Generation under Complex Interactions
by: Liang, Han, et al.
Published: (2023)
by: Liang, Han, et al.
Published: (2023)
UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?
by: Liu, Yuanxin, et al.
Published: (2025)
by: Liu, Yuanxin, et al.
Published: (2025)
3D Spatial Understanding in MLLMs: Disambiguation and Evaluation
by: Chang, Chun-Peng, et al.
Published: (2024)
by: Chang, Chun-Peng, et al.
Published: (2024)
Similar Items
-
EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy
by: Li, Jinzhao, et al.
Published: (2026) -
PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers
by: Lin, Yuchen, et al.
Published: (2025) -
MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second
by: Lin, Chenguo, et al.
Published: (2025) -
X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding
by: Sun, Peiwen, et al.
Published: (2026) -
MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning
by: Wang, Yueqian, et al.
Published: (2025)