StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Haolin, Tang, Feilong, Zhao, Lingxiao, Zhuang, Xinlin, Lu, Yifan, An, Xiang, Hu, Ming, Zhang, Xiaofeng, Swikir, Abdalla, He, Junjun, Ge, Zongyuan, Khan, Muhammad Haris, Razzak, Imran |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Efficient Medical Reasoning with Minimal Fine-Tuning Data
by: Zhuang, Xinlin, et al.
Published: (2025)
by: Zhuang, Xinlin, et al.
Published: (2025)
ScalingNoise: Scaling Inference-Time Search for Generating Infinite Videos
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
Finite Abstractions of Network of Impulsive Systems Using Dissipativity Approach
by: Swikir, Abdalla
Published: (2024)
by: Swikir, Abdalla
Published: (2024)
Towards Dynamic 3D Reconstruction of Hand-Instrument Interaction in Ophthalmic Surgery
by: Hu, Ming, et al.
Published: (2025)
by: Hu, Ming, et al.
Published: (2025)
Robust Multimodal Learning for Ophthalmic Disease Grading via Disentangled Representation
by: Wang, Xinkun, et al.
Published: (2025)
by: Wang, Xinkun, et al.
Published: (2025)
Towards Robust Visual Continual Learning with Multi-Prototype Supervision
by: Liu, Xiwei, et al.
Published: (2025)
by: Liu, Xiwei, et al.
Published: (2025)
StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
by: Wang, Junxi, et al.
Published: (2026)
by: Wang, Junxi, et al.
Published: (2026)
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning
by: Deria, Ankan, et al.
Published: (2025)
by: Deria, Ankan, et al.
Published: (2025)
Phenome-Wide Multi-Omics Integration Uncovers Distinct Archetypes of Human Aging
by: Li, Huifa, et al.
Published: (2025)
by: Li, Huifa, et al.
Published: (2025)
Discriminating retinal microvascular and neuronal differences related to migraines: Deep Learning based Crossectional Study
by: Tang, Feilong, et al.
Published: (2024)
by: Tang, Feilong, et al.
Published: (2024)
ClinCoT: Clinical-Aware Visual Chain-of-Thought for Medical Vision Language Models
by: Liu, Xiwei, et al.
Published: (2026)
by: Liu, Xiwei, et al.
Published: (2026)
Anticipatory and Adaptive Footstep Streaming for Teleoperated Bipedal Robots
by: Penco, Luigi, et al.
Published: (2025)
by: Penco, Luigi, et al.
Published: (2025)
PG-SAM: Prior-Guided SAM with Medical for Multi-organ Segmentation
by: Zhong, Yiheng, et al.
Published: (2025)
by: Zhong, Yiheng, et al.
Published: (2025)
DeLo: Dual Decomposed Low-Rank Experts Collaboration for Continual Missing Modality Learning
by: Liu, Xiwei, et al.
Published: (2026)
by: Liu, Xiwei, et al.
Published: (2026)
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
by: Tang, Feilong, et al.
Published: (2025)
by: Tang, Feilong, et al.
Published: (2025)
scAGC: Learning Adaptive Cell Graphs with Contrastive Guidance for Single-Cell Clustering
by: Li, Huifa, et al.
Published: (2025)
by: Li, Huifa, et al.
Published: (2025)
MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World Conversation
by: Xue, Haochen, et al.
Published: (2025)
by: Xue, Haochen, et al.
Published: (2025)
PEARL: Personalized Streaming Video Understanding Model
by: Zheng, Yuanhong, et al.
Published: (2026)
by: Zheng, Yuanhong, et al.
Published: (2026)
LATA: Laplacian-Assisted Transductive Adaptation for Conformal Uncertainty in Medical VLMs
by: Bozorgtabar, Behzad, et al.
Published: (2026)
by: Bozorgtabar, Behzad, et al.
Published: (2026)
An Experimental Study of Low-Latency Video Streaming over 5G
by: Khan, Imran, et al.
Published: (2024)
by: Khan, Imran, et al.
Published: (2024)
StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
by: Wang, Yifei, et al.
Published: (2025)
by: Wang, Yifei, et al.
Published: (2025)
StreamingTOM: Streaming Token Compression for Efficient Video Understanding
by: Chen, Xueyi, et al.
Published: (2025)
by: Chen, Xueyi, et al.
Published: (2025)
StreamingVLM: Real-Time Understanding for Infinite Video Streams
by: Xu, Ruyi, et al.
Published: (2025)
by: Xu, Ruyi, et al.
Published: (2025)
DermAgent: A Self-Reflective Agentic System for Dermatological Image Analysis with Multi-Tool Reasoning and Traceable Decision-Making
by: Liu, Yize, et al.
Published: (2026)
by: Liu, Yize, et al.
Published: (2026)
StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
by: Xie, Ming, et al.
Published: (2026)
by: Xie, Ming, et al.
Published: (2026)
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
by: Lin, Junming, et al.
Published: (2024)
by: Lin, Junming, et al.
Published: (2024)
Anticipatory Planning for Multimodal AI Agents
by: Liang, Yongyuan, et al.
Published: (2026)
by: Liang, Yongyuan, et al.
Published: (2026)
Revisable by Design: A Theory of Streaming LLM Agent Execution
by: Zhai, Zhiyuan, et al.
Published: (2026)
by: Zhai, Zhiyuan, et al.
Published: (2026)
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding
by: Wu, Hang, et al.
Published: (2026)
by: Wu, Hang, et al.
Published: (2026)
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
by: Lee, Daeun, et al.
Published: (2025)
by: Lee, Daeun, et al.
Published: (2025)
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
by: Yang, Yanlai, et al.
Published: (2025)
by: Yang, Yanlai, et al.
Published: (2025)
Generative AI Agents for Controllable and Protected Content Creation
by: Khan, Haris, et al.
Published: (2026)
by: Khan, Haris, et al.
Published: (2026)
OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents
by: Shabbir, Akashah, et al.
Published: (2026)
by: Shabbir, Akashah, et al.
Published: (2026)
StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
by: Chen, Yilong, et al.
Published: (2025)
by: Chen, Yilong, et al.
Published: (2025)
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling
by: Luo, Yawen, et al.
Published: (2026)
by: Luo, Yawen, et al.
Published: (2026)
FairStream: Fair Multimedia Streaming Benchmark for Reinforcement Learning Agents
by: Weil, Jannis, et al.
Published: (2024)
by: Weil, Jannis, et al.
Published: (2024)
Joint Optimization of Streaming and Non-Streaming Automatic Speech Recognition with Multi-Decoder and Knowledge Distillation
by: Shakeel, Muhammad, et al.
Published: (2024)
by: Shakeel, Muhammad, et al.
Published: (2024)
Video Streaming with Kairos: An MPC-Based ABR with Streaming-Aware Throughput Prediction
by: Zhong, Ziyu, et al.
Published: (2025)
by: Zhong, Ziyu, et al.
Published: (2025)
Devil's Advocate: Anticipatory Reflection for LLM Agents
by: Wang, Haoyu, et al.
Published: (2024)
by: Wang, Haoyu, et al.
Published: (2024)
A Simple Baseline for Streaming Video Understanding
by: Shen, Yujiao, et al.
Published: (2026)
by: Shen, Yujiao, et al.
Published: (2026)
Similar Items
-
Towards Efficient Medical Reasoning with Minimal Fine-Tuning Data
by: Zhuang, Xinlin, et al.
Published: (2025) -
ScalingNoise: Scaling Inference-Time Search for Generating Infinite Videos
by: Yang, Haolin, et al.
Published: (2025) -
Finite Abstractions of Network of Impulsive Systems Using Dissipativity Approach
by: Swikir, Abdalla
Published: (2024) -
Towards Dynamic 3D Reconstruction of Hand-Instrument Interaction in Ophthalmic Surgery
by: Hu, Ming, et al.
Published: (2025) -
Robust Multimodal Learning for Ophthalmic Disease Grading via Disentangled Representation
by: Wang, Xinkun, et al.
Published: (2025)