Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Ke, Tang, Jiaqi, Guo, Bin, Han, Xueting, Xu, Ruonan, He, Qingfeng, Wang, Ziheng, Wang, Xu, Chen, Qifeng, Yu, Zhiwen, Liu, Yunhao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding
by: Guo, Pengze, et al.
Published: (2026)
by: Guo, Pengze, et al.
Published: (2026)
Hawk: Learning to Understand Open-World Video Anomalies
by: Tang, Jiaqi, et al.
Published: (2024)
by: Tang, Jiaqi, et al.
Published: (2024)
StreamingVLM: Real-Time Understanding for Infinite Video Streams
by: Xu, Ruyi, et al.
Published: (2025)
by: Xu, Ruyi, et al.
Published: (2025)
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
by: Lee, Daeun, et al.
Published: (2025)
by: Lee, Daeun, et al.
Published: (2025)
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
by: Zhao, Ruixiang, et al.
Published: (2026)
by: Zhao, Ruixiang, et al.
Published: (2026)
Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding
by: Zheng, Yikai, et al.
Published: (2026)
by: Zheng, Yikai, et al.
Published: (2026)
From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench
by: Xu, Ke, et al.
Published: (2026)
by: Xu, Ke, et al.
Published: (2026)
PaSBench-Video: A Streaming Video Benchmark for Proactive Safety Warning
by: Zhao, Yusong, et al.
Published: (2026)
by: Zhao, Yusong, et al.
Published: (2026)
AdaShadow: Responsive Test-time Model Adaptation in Non-stationary Mobile Environments
by: Fang, Cheng, et al.
Published: (2024)
by: Fang, Cheng, et al.
Published: (2024)
Explicit Topology Optimization Based on the Joint‐Driven Moving Morphable Components
by: Jiaqi Xu, et al.
Published: (2025)
by: Jiaqi Xu, et al.
Published: (2025)
Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video
by: Zhang, Yulin, et al.
Published: (2025)
by: Zhang, Yulin, et al.
Published: (2025)
StreamPro: From Reactive Perception to Proactive Decision-Making in Streaming Video
by: Li, Ao, et al.
Published: (2026)
by: Li, Ao, et al.
Published: (2026)
Proactive Memory for Ad-Hoc Recall over Streaming Dialogues
by: Wang, Bingbing, et al.
Published: (2026)
by: Wang, Bingbing, et al.
Published: (2026)
EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams
by: Ran, Dongchuan, et al.
Published: (2026)
by: Ran, Dongchuan, et al.
Published: (2026)
StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant
by: Wang, Haibo, et al.
Published: (2025)
by: Wang, Haibo, et al.
Published: (2025)
FedRC: A Rapid-Converged Hierarchical Federated Learning Framework in Street Scene Semantic Understanding
by: Kou, Wei-Bin, et al.
Published: (2024)
by: Kou, Wei-Bin, et al.
Published: (2024)
Distill, Diffuse, and Semanticize (DDS): Annotation-Free 3D Scene Understanding Based on Multi-Granularity Distillation and Graph-Diffusion-Based Segmentation
by: Wang, Yijing, et al.
Published: (2026)
by: Wang, Yijing, et al.
Published: (2026)
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
by: Qi, Zhangyang, et al.
Published: (2025)
by: Qi, Zhangyang, et al.
Published: (2025)
Streaming Long Video Understanding with Large Language Models
by: Qian, Rui, et al.
Published: (2024)
by: Qian, Rui, et al.
Published: (2024)
StreamingEval: A Unified Evaluation Protocol towards Realistic Streaming Video Understanding
by: Tang, Guowei, et al.
Published: (2026)
by: Tang, Guowei, et al.
Published: (2026)
StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
by: Xie, Ming, et al.
Published: (2026)
by: Xie, Ming, et al.
Published: (2026)
StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
by: Wang, Yifei, et al.
Published: (2025)
by: Wang, Yifei, et al.
Published: (2025)
Enhancing Large Vision Model in Street Scene Semantic Understanding through Leveraging Posterior Optimization Trajectory
by: Kou, Wei-Bin, et al.
Published: (2025)
by: Kou, Wei-Bin, et al.
Published: (2025)
LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
by: Yang, Zhenyu, et al.
Published: (2025)
by: Yang, Zhenyu, et al.
Published: (2025)
StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
Difference-in-Differences with Interference
by: Xu, Ruonan
Published: (2023)
by: Xu, Ruonan
Published: (2023)
Open-Vocabulary Octree-Graph for 3D Scene Understanding
by: Wang, Zhigang, et al.
Published: (2024)
by: Wang, Zhigang, et al.
Published: (2024)
Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?
by: Xu, Boshen, et al.
Published: (2024)
by: Xu, Boshen, et al.
Published: (2024)
EgoGraph: Temporal Knowledge Graph for Egocentric Video Understanding
by: Sun, Shitong, et al.
Published: (2026)
by: Sun, Shitong, et al.
Published: (2026)
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
<SOG_k>: One LLM Token for Explicit Graph Structural Understanding
by: Wu, Jingyao, et al.
Published: (2026)
by: Wu, Jingyao, et al.
Published: (2026)
Detecting Fake Reviewer Groups in Dynamic Networks: An Adaptive Graph Learning Method
by: Zhang, Jing, et al.
Published: (2026)
by: Zhang, Jing, et al.
Published: (2026)
StreamingTOM: Streaming Token Compression for Efficient Video Understanding
by: Chen, Xueyi, et al.
Published: (2025)
by: Chen, Xueyi, et al.
Published: (2025)
An Efficient Streaming Video Understanding Framework with Agentic Control
by: Liu, Jinming, et al.
Published: (2026)
by: Liu, Jinming, et al.
Published: (2026)
Traffic Regulation-aware Path Planning with Regulation Databases and Vision-Language Models
by: Han, Xu, et al.
Published: (2025)
by: Han, Xu, et al.
Published: (2025)
StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
by: Wang, Junxi, et al.
Published: (2026)
by: Wang, Junxi, et al.
Published: (2026)
FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding
by: Xie, Yiweng, et al.
Published: (2026)
by: Xie, Yiweng, et al.
Published: (2026)
GSStream: 3D Gaussian Splatting based Volumetric Scene Streaming System
by: Tang, Zhiye, et al.
Published: (2026)
by: Tang, Zhiye, et al.
Published: (2026)
Dynamic Gaussian Scene Reconstruction from Unsynchronized Videos
by: Xu, Zhixin, et al.
Published: (2025)
by: Xu, Zhixin, et al.
Published: (2025)
VL-KnG: Persistent Spatiotemporal Knowledge Graphs from Egocentric Video for Embodied Scene Understanding
by: Mdfaa, Mohamad Al, et al.
Published: (2025)
by: Mdfaa, Mohamad Al, et al.
Published: (2025)
Similar Items
-
EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding
by: Guo, Pengze, et al.
Published: (2026) -
Hawk: Learning to Understand Open-World Video Anomalies
by: Tang, Jiaqi, et al.
Published: (2024) -
StreamingVLM: Real-Time Understanding for Infinite Video Streams
by: Xu, Ruyi, et al.
Published: (2025) -
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
by: Lee, Daeun, et al.
Published: (2025) -
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
by: Zhao, Ruixiang, et al.
Published: (2026)