OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Yibin, Xu, Jilan, Di, Shangzhe, Wu, Haoning, Xie, Weidi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Streaming Video Representation via Multitask Training
by: Yan, Yibin, et al.
Published: (2025)
by: Yan, Yibin, et al.
Published: (2025)
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023)
by: Di, Shangzhe, et al.
Published: (2023)
Revisiting Multi-Task Visual Representation Learning
by: Di, Shangzhe, et al.
Published: (2026)
by: Di, Shangzhe, et al.
Published: (2026)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
by: Chen, Qirui, et al.
Published: (2024)
by: Chen, Qirui, et al.
Published: (2024)
EchoSight: Advancing Visual-Language Models with Wiki Knowledge
by: Yan, Yibin, et al.
Published: (2024)
by: Yan, Yibin, et al.
Published: (2024)
SoccerMaster: A Vision Foundation Model for Soccer Understanding
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation
by: Shi, Yudi, et al.
Published: (2024)
by: Shi, Yudi, et al.
Published: (2024)
Count Anything at Any Granularity
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
by: Xie, Ming, et al.
Published: (2026)
by: Xie, Ming, et al.
Published: (2026)
Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
by: Li, Zeqian, et al.
Published: (2025)
by: Li, Zeqian, et al.
Published: (2025)
Action Emergence from Streaming Intent
by: Jing, Pengfei, et al.
Published: (2026)
by: Jing, Pengfei, et al.
Published: (2026)
SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass
by: Meng, Yanxu, et al.
Published: (2025)
by: Meng, Yanxu, et al.
Published: (2025)
MatchTime: Towards Automatic Soccer Game Commentary Generation
by: Rao, Jiayuan, et al.
Published: (2024)
by: Rao, Jiayuan, et al.
Published: (2024)
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
by: Zhao, Ruixiang, et al.
Published: (2026)
by: Zhao, Ruixiang, et al.
Published: (2026)
Instant Gaussian Stream: Fast and Generalizable Streaming of Dynamic Scene Reconstruction via Gaussian Splatting
by: Yan, Jinbo, et al.
Published: (2025)
by: Yan, Jinbo, et al.
Published: (2025)
StreamGS: Online Generalizable Gaussian Splatting Reconstruction for Unposed Image Streams
by: LI, Yang, et al.
Published: (2025)
by: LI, Yang, et al.
Published: (2025)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
by: Di, Shangzhe, et al.
Published: (2025)
by: Di, Shangzhe, et al.
Published: (2025)
ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding
by: Tian, Xueyun, et al.
Published: (2026)
by: Tian, Xueyun, et al.
Published: (2026)
Towards Universal Soccer Video Understanding
by: Rao, Jiayuan, et al.
Published: (2024)
by: Rao, Jiayuan, et al.
Published: (2024)
Intelligent Grimm -- Open-ended Visual Storytelling via Latent Diffusion Models
by: Liu, Chang, et al.
Published: (2023)
by: Liu, Chang, et al.
Published: (2023)
Multi-Agent System for Comprehensive Soccer Understanding
by: Rao, Jiayuan, et al.
Published: (2025)
by: Rao, Jiayuan, et al.
Published: (2025)
Retrieval-Augmented Egocentric Video Captioning
by: Xu, Jilan, et al.
Published: (2024)
by: Xu, Jilan, et al.
Published: (2024)
StreamSplat: Towards Online Dynamic 3D Reconstruction from Uncalibrated Video Streams
by: Wu, Zike, et al.
Published: (2025)
by: Wu, Zike, et al.
Published: (2025)
OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts
by: Wang, Yuxuan, et al.
Published: (2025)
by: Wang, Yuxuan, et al.
Published: (2025)
HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction
by: Cheng, Chong, et al.
Published: (2026)
by: Cheng, Chong, et al.
Published: (2026)
MRGen: Segmentation Data Engine for Underrepresented MRI Modalities
by: Wu, Haoning, et al.
Published: (2024)
by: Wu, Haoning, et al.
Published: (2024)
Dual-Stream Alignment for Action Segmentation
by: Gammulle, Harshala, et al.
Published: (2025)
by: Gammulle, Harshala, et al.
Published: (2025)
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
by: Shi, Yudi, et al.
Published: (2026)
by: Shi, Yudi, et al.
Published: (2026)
OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation
by: Song, Yiren, et al.
Published: (2026)
by: Song, Yiren, et al.
Published: (2026)
MeMix: Writing Less, Remembering More for Streaming 3D Reconstruction
by: Dong, Jiacheng, et al.
Published: (2026)
by: Dong, Jiacheng, et al.
Published: (2026)
StreamingVLA: Streaming Vision-Language-Action Model with Action Flow Matching and Adaptive Early Observation
by: Shi, Yiran, et al.
Published: (2026)
by: Shi, Yiran, et al.
Published: (2026)
StreamPro: From Reactive Perception to Proactive Decision-Making in Streaming Video
by: Li, Ao, et al.
Published: (2026)
by: Li, Ao, et al.
Published: (2026)
QUEST: Query Stream for Practical Cooperative Perception
by: Fan, Siqi, et al.
Published: (2023)
by: Fan, Siqi, et al.
Published: (2023)
ActionPlan: Future-Aware Streaming Motion Synthesis via Frame-Level Action Planning
by: Nazarenus, Eric, et al.
Published: (2026)
by: Nazarenus, Eric, et al.
Published: (2026)
D-SCo: Dual-Stream Conditional Diffusion for Monocular Hand-Held Object Reconstruction
by: Fu, Bowen, et al.
Published: (2023)
by: Fu, Bowen, et al.
Published: (2023)
Geometric Context Transformer for Streaming 3D Reconstruction
by: Chen, Lin-Zhuo, et al.
Published: (2026)
by: Chen, Lin-Zhuo, et al.
Published: (2026)
Progressive Cross-Stream Cooperation in Spatial and Temporal Domain for Action Localization
by: Su, Rui, et al.
Published: (2019)
by: Su, Rui, et al.
Published: (2019)
Complementing Event Streams and RGB Frames for Hand Mesh Reconstruction
by: Jiang, Jianping, et al.
Published: (2024)
by: Jiang, Jianping, et al.
Published: (2024)
ActionSwitch: Class-agnostic Detection of Simultaneous Actions in Streaming Videos
by: Kang, Hyolim, et al.
Published: (2024)
by: Kang, Hyolim, et al.
Published: (2024)
SLARM: Streaming and Language-Aligned Reconstruction Model for Dynamic Scenes
by: Qiu, Zhicheng, et al.
Published: (2026)
by: Qiu, Zhicheng, et al.
Published: (2026)
Similar Items
-
Learning Streaming Video Representation via Multitask Training
by: Yan, Yibin, et al.
Published: (2025) -
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023) -
Revisiting Multi-Task Visual Representation Learning
by: Di, Shangzhe, et al.
Published: (2026) -
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
by: Chen, Qirui, et al.
Published: (2024) -
EchoSight: Advancing Visual-Language Models with Wiki Knowledge
by: Yan, Yibin, et al.
Published: (2024)