Making Every Frame Matter: Continuous Activity Recognition in Streaming Video via Adaptive Video Context Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Hao, Bai, Donglin, Jiang, Shiqi, Zhang, Qianxi, Yang, Yifan, Ding, Xin, Cao, Ting, Liu, Yunxin, Xu, Fengyuan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition
by: Ding, Xin, et al.
Published: (2025)
by: Ding, Xin, et al.
Published: (2025)
Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding
by: Zheng, Yikai, et al.
Published: (2026)
by: Zheng, Yikai, et al.
Published: (2026)
AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation
by: Ding, Xin, et al.
Published: (2025)
by: Ding, Xin, et al.
Published: (2025)
Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
by: Wang, Chendong, et al.
Published: (2025)
by: Wang, Chendong, et al.
Published: (2025)
MemCompiler: Compile, Don't Inject -- State-Conditioned Memory for Embodied Agents
by: Ding, Xin, et al.
Published: (2026)
by: Ding, Xin, et al.
Published: (2026)
Scaling Up On-Device LLMs via Active-Weight Swapping Between DRAM and Flash
by: Jia, Fucheng, et al.
Published: (2025)
by: Jia, Fucheng, et al.
Published: (2025)
EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents
by: Ju, Ruofei, et al.
Published: (2026)
by: Ju, Ruofei, et al.
Published: (2026)
VideoScan: Enabling Efficient Streaming Video Understanding via Frame-level Semantic Carriers
by: Li, Ruanjun, et al.
Published: (2025)
by: Li, Ruanjun, et al.
Published: (2025)
AVA: Towards Agentic Video Analytics with Vision Language Models
by: Yan, Yuxuan, et al.
Published: (2025)
by: Yan, Yuxuan, et al.
Published: (2025)
FC-VFI: Faithful and Consistent Video Frame Interpolation for High-FPS Slow Motion Video Generation
by: Ding, Ganggui, et al.
Published: (2026)
by: Ding, Ganggui, et al.
Published: (2026)
Efficient Remote KV Cache Reuse with GPU-native Video Codec
by: Mi, Liang, et al.
Published: (2026)
by: Mi, Liang, et al.
Published: (2026)
StreamOptix: A Cross-layer Adaptive Video Delivery Scheme
by: Liu, Mufan, et al.
Published: (2024)
by: Liu, Mufan, et al.
Published: (2024)
End-to-End Dense Video Grounding via Parallel Regression
by: Shi, Fengyuan, et al.
Published: (2021)
by: Shi, Fengyuan, et al.
Published: (2021)
BiSwift: Bandwidth Orchestrator for Multi-Stream Video Analytics on Edge
by: Sun, Lin, et al.
Published: (2023)
by: Sun, Lin, et al.
Published: (2023)
Adaptive 3D Gaussian Splatting Video Streaming
by: Gong, Han, et al.
Published: (2025)
by: Gong, Han, et al.
Published: (2025)
CogStream: Context-guided Streaming Video Question Answering
by: Zhao, Zicheng, et al.
Published: (2025)
by: Zhao, Zicheng, et al.
Published: (2025)
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
by: Zhang, Peng, et al.
Published: (2026)
by: Zhang, Peng, et al.
Published: (2026)
Sequence-Adaptive Video Prediction in Continuous Streams using Diffusion Noise Optimization
by: Azar, Sina Mokhtarzadeh, et al.
Published: (2025)
by: Azar, Sina Mokhtarzadeh, et al.
Published: (2025)
Beyond the Last Frame: Process-aware Evaluation for Generative Video Reasoning
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
Video Frame Interpolation for Polarization via Swin-Transformer
by: Huang, Feng, et al.
Published: (2024)
by: Huang, Feng, et al.
Published: (2024)
On Denoising Walking Videos for Gait Recognition
by: Jin, Dongyang, et al.
Published: (2025)
by: Jin, Dongyang, et al.
Published: (2025)
Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
by: Ling Team, et al.
Published: (2025)
by: Ling Team, et al.
Published: (2025)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
by: Di, Shangzhe, et al.
Published: (2025)
by: Di, Shangzhe, et al.
Published: (2025)
Event-based Continuous Color Video Decompression from Single Frames
by: Wang, Ziyun, et al.
Published: (2023)
by: Wang, Ziyun, et al.
Published: (2023)
DeformStream: Deformation-based Adaptive Volumetric Video Streaming
by: Li, Boyan, et al.
Published: (2024)
by: Li, Boyan, et al.
Published: (2024)
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing
by: Lee, Hosu, et al.
Published: (2024)
by: Lee, Hosu, et al.
Published: (2024)
Online Misinformation Detection in Live Streaming Videos
by: Cao, Rui
Published: (2025)
by: Cao, Rui
Published: (2025)
StreamPro: From Reactive Perception to Proactive Decision-Making in Streaming Video
by: Li, Ao, et al.
Published: (2026)
by: Li, Ao, et al.
Published: (2026)
Learning from One Continuous Video Stream
by: Carreira, João, et al.
Published: (2023)
by: Carreira, João, et al.
Published: (2023)
Characterizing User Platforms for Video Streaming in Broadband Networks
by: Wang, Yifan, et al.
Published: (2024)
by: Wang, Yifan, et al.
Published: (2024)
Adaptive Greedy Frame Selection for Long Video Understanding
by: Huang, Yuning, et al.
Published: (2026)
by: Huang, Yuning, et al.
Published: (2026)
Selective Volume Mixup for Video Action Recognition
by: Tan, Yi, et al.
Published: (2023)
by: Tan, Yi, et al.
Published: (2023)
Sculptor: Empowering LLMs with Cognitive Agency via Active Context Management
by: Li, Mo, et al.
Published: (2025)
by: Li, Mo, et al.
Published: (2025)
Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
by: Zhang, Lvmin, et al.
Published: (2025)
by: Zhang, Lvmin, et al.
Published: (2025)
Long-Context Autoregressive Video Modeling with Next-Frame Prediction
by: Gu, Yuchao, et al.
Published: (2025)
by: Gu, Yuchao, et al.
Published: (2025)
A Gift for Every Teacher in Every Language: The Video Library of Classroom Practices
by: Duncan, Greg
Published: (2008)
by: Duncan, Greg
Published: (2008)
Perception-Oriented Video Frame Interpolation via Asymmetric Blending
by: Wu, Guangyang, et al.
Published: (2024)
by: Wu, Guangyang, et al.
Published: (2024)
WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs
by: Zhang, Yulin, et al.
Published: (2026)
by: Zhang, Yulin, et al.
Published: (2026)
Progressive Frame Patching for FoV-based Point Cloud Video Streaming
by: Zong, Tongyu, et al.
Published: (2023)
by: Zong, Tongyu, et al.
Published: (2023)
LFS: Learnable Frame Selector for Event-Aware and Temporally Diverse Video Captioning
by: Chao, Lianying, et al.
Published: (2026)
by: Chao, Lianying, et al.
Published: (2026)
Similar Items
-
StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition
by: Ding, Xin, et al.
Published: (2025) -
Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding
by: Zheng, Yikai, et al.
Published: (2026) -
AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation
by: Ding, Xin, et al.
Published: (2025) -
Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
by: Wang, Chendong, et al.
Published: (2025) -
MemCompiler: Compile, Don't Inject -- State-Conditioned Memory for Embodied Agents
by: Ding, Xin, et al.
Published: (2026)