Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Chatterjee, Dibyadip, Remelli, Edoardo, Song, Yale, Tekin, Bugra, Mittal, Abhay, Bhatnagar, Bharat, Camgöz, Necati Cihan, Hampali, Shreyas, Sauser, Eric, Ma, Shugao, Yao, Angela, Sener, Fadime |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DiffH2O: Diffusion-Based Synthesis of Hand-Object Interactions from Textual Descriptions
by: Christen, Sammy, et al.
Published: (2024)
by: Christen, Sammy, et al.
Published: (2024)
X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization
by: Kukleva, Anna, et al.
Published: (2024)
by: Kukleva, Anna, et al.
Published: (2024)
PALM: A Dataset and Baseline for Learning Multi-subject Hand Prior
by: Fan, Zicong, et al.
Published: (2025)
by: Fan, Zicong, et al.
Published: (2025)
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
by: Pang, Zhanzhong, et al.
Published: (2026)
by: Pang, Zhanzhong, et al.
Published: (2026)
On the Utility of 3D Hand Poses for Action Recognition
by: Shamil, Md Salman, et al.
Published: (2024)
by: Shamil, Md Salman, et al.
Published: (2024)
Don't Pause! Every prediction matters in a streaming video
by: Chatterjee, Dibyadip, et al.
Published: (2026)
by: Chatterjee, Dibyadip, et al.
Published: (2026)
On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action Understanding
by: Pang, Zhanzhong, et al.
Published: (2026)
by: Pang, Zhanzhong, et al.
Published: (2026)
Sign2GPT: Leveraging Large Language Models for Gloss-Free Sign Language Translation
by: Wong, Ryan, et al.
Published: (2024)
by: Wong, Ryan, et al.
Published: (2024)
SignRep: Enhancing Self-Supervised Sign Representations
by: Wong, Ryan, et al.
Published: (2025)
by: Wong, Ryan, et al.
Published: (2025)
Towards Privacy-Aware Sign Language Translation at Scale
by: Rust, Phillip, et al.
Published: (2024)
by: Rust, Phillip, et al.
Published: (2024)
VideoLLM-online: Online Video Large Language Model for Streaming Video
by: Chen, Joya, et al.
Published: (2024)
by: Chen, Joya, et al.
Published: (2024)
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
by: Guan, Yiran, et al.
Published: (2026)
by: Guan, Yiran, et al.
Published: (2026)
HuMoCon: Concept Discovery for Human Motion Understanding
by: Fang, Qihang, et al.
Published: (2025)
by: Fang, Qihang, et al.
Published: (2025)
SneakPeek: Future-Guided Instructional Streaming Video Generation
by: Hong, Cheeun, et al.
Published: (2025)
by: Hong, Cheeun, et al.
Published: (2025)
POET: Prompt Offset Tuning for Continual Human Action Adaptation
by: Garg, Prachi, et al.
Published: (2025)
by: Garg, Prachi, et al.
Published: (2025)
VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation
by: Wu, Shiwei, et al.
Published: (2024)
by: Wu, Shiwei, et al.
Published: (2024)
WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs
by: Zhang, Yulin, et al.
Published: (2026)
by: Zhang, Yulin, et al.
Published: (2026)
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
by: Qin, Jialong, et al.
Published: (2025)
by: Qin, Jialong, et al.
Published: (2025)
VideoLLM Benchmarks and Evaluation: A Survey
by: Kumar, Yogesh
Published: (2025)
by: Kumar, Yogesh
Published: (2025)
Context-Enhanced Memory-Refined Transformer for Online Action Detection
by: Pang, Zhanzhong, et al.
Published: (2025)
by: Pang, Zhanzhong, et al.
Published: (2025)
Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
by: Chung, Hyungjin, et al.
Published: (2025)
by: Chung, Hyungjin, et al.
Published: (2025)
Geometry-Guided Camera Motion Understanding in VideoLLMs
by: Feng, Haoan, et al.
Published: (2026)
by: Feng, Haoan, et al.
Published: (2026)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
by: Li, Chenglin, et al.
Published: (2026)
by: Li, Chenglin, et al.
Published: (2026)
Lost in Time: A New Temporal Benchmark for VideoLLMs
by: Cores, Daniel, et al.
Published: (2024)
by: Cores, Daniel, et al.
Published: (2024)
Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
by: Kim, Minji, et al.
Published: (2025)
by: Kim, Minji, et al.
Published: (2025)
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
by: Liang, Yujia, et al.
Published: (2025)
by: Liang, Yujia, et al.
Published: (2025)
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding
by: Fang, Pengcheng, et al.
Published: (2025)
by: Fang, Pengcheng, et al.
Published: (2025)
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
by: Wang, Haibo, et al.
Published: (2024)
by: Wang, Haibo, et al.
Published: (2024)
Cost-Sensitive Learning for Long-Tailed Temporal Action Segmentation
by: Pang, Zhanzhong, et al.
Published: (2025)
by: Pang, Zhanzhong, et al.
Published: (2025)
Long-Tail Temporal Action Segmentation with Group-wise Temporal Logit Adjustment
by: Pang, Zhanzhong, et al.
Published: (2024)
by: Pang, Zhanzhong, et al.
Published: (2024)
Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency
by: Li, Hongyu, et al.
Published: (2025)
by: Li, Hongyu, et al.
Published: (2025)
Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
by: Kumar, Yogesh
Published: (2025)
by: Kumar, Yogesh
Published: (2025)
EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs
by: Li, Jiameng, et al.
Published: (2026)
by: Li, Jiameng, et al.
Published: (2026)
Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
by: Cai, Jianfeng, et al.
Published: (2025)
by: Cai, Jianfeng, et al.
Published: (2025)
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
Proact-VL: A Proactive VideoLLM for Real-Time AI Companions
by: Yan, Weicai, et al.
Published: (2026)
by: Yan, Weicai, et al.
Published: (2026)
Procedural Game Level Design with Deep Reinforcement Learning
by: Özkan, Miraç Buğra
Published: (2025)
by: Özkan, Miraç Buğra
Published: (2025)
Reseña de "Projeto Conservação Preventiva em Bibliotecas e Arquivos - CPBA"
by: Alda L. Sauser
Published: (2005)
by: Alda L. Sauser
Published: (2005)
Template Free Reconstruction of Human-object Interaction with Procedural Interaction Generation
by: Xie, Xianghui, et al.
Published: (2023)
by: Xie, Xianghui, et al.
Published: (2023)
Similar Items
-
DiffH2O: Diffusion-Based Synthesis of Hand-Object Interactions from Textual Descriptions
by: Christen, Sammy, et al.
Published: (2024) -
X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization
by: Kukleva, Anna, et al.
Published: (2024) -
PALM: A Dataset and Baseline for Learning Multi-subject Hand Prior
by: Fan, Zicong, et al.
Published: (2025) -
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
by: Pang, Zhanzhong, et al.
Published: (2026) -
On the Utility of 3D Hand Poses for Action Recognition
by: Shamil, Md Salman, et al.
Published: (2024)