Don't Pause! Every prediction matters in a streaming video
Fuente:
arXiv
Saved in:
| Main Authors: | Chatterjee, Dibyadip, Pang, Zhanzhong, Sener, Fadime, Song, Yale, Yao, Angela |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
by: Pang, Zhanzhong, et al.
Published: (2026)
by: Pang, Zhanzhong, et al.
Published: (2026)
On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action Understanding
by: Pang, Zhanzhong, et al.
Published: (2026)
by: Pang, Zhanzhong, et al.
Published: (2026)
Context-Enhanced Memory-Refined Transformer for Online Action Detection
by: Pang, Zhanzhong, et al.
Published: (2025)
by: Pang, Zhanzhong, et al.
Published: (2025)
Cost-Sensitive Learning for Long-Tailed Temporal Action Segmentation
by: Pang, Zhanzhong, et al.
Published: (2025)
by: Pang, Zhanzhong, et al.
Published: (2025)
Long-Tail Temporal Action Segmentation with Group-wise Temporal Logit Adjustment
by: Pang, Zhanzhong, et al.
Published: (2024)
by: Pang, Zhanzhong, et al.
Published: (2024)
On the Utility of 3D Hand Poses for Action Recognition
by: Shamil, Md Salman, et al.
Published: (2024)
by: Shamil, Md Salman, et al.
Published: (2024)
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
by: Chatterjee, Dibyadip, et al.
Published: (2025)
by: Chatterjee, Dibyadip, et al.
Published: (2025)
X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization
by: Kukleva, Anna, et al.
Published: (2024)
by: Kukleva, Anna, et al.
Published: (2024)
NeIn: Telling What You Don't Want
by: Bui, Nhat-Tan, et al.
Published: (2024)
by: Bui, Nhat-Tan, et al.
Published: (2024)
PALM: A Dataset and Baseline for Learning Multi-subject Hand Prior
by: Fan, Zicong, et al.
Published: (2025)
by: Fan, Zicong, et al.
Published: (2025)
Show, Don't Tell: Morphing Latent Reasoning into Image Generation
by: Chen, Harold Haodong, et al.
Published: (2026)
by: Chen, Harold Haodong, et al.
Published: (2026)
Don't Look into the Dark: Latent Codes for Pluralistic Image Inpainting
by: Chen, Haiwei, et al.
Published: (2024)
by: Chen, Haiwei, et al.
Published: (2024)
Visually Dehallucinative Instruction Generation: Know What You Don't Know
by: Cha, Sungguk, et al.
Published: (2024)
by: Cha, Sungguk, et al.
Published: (2024)
ProCreate, Don't Reproduce! Propulsive Energy Diffusion for Creative Generation
by: Lu, Jack, et al.
Published: (2024)
by: Lu, Jack, et al.
Published: (2024)
Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent
by: Ci, En, et al.
Published: (2025)
by: Ci, En, et al.
Published: (2025)
SneakPeek: Future-Guided Instructional Streaming Video Generation
by: Hong, Cheeun, et al.
Published: (2025)
by: Hong, Cheeun, et al.
Published: (2025)
Vision Transformers Don't Need Trained Registers
by: Jiang, Nick, et al.
Published: (2025)
by: Jiang, Nick, et al.
Published: (2025)
Don't let the information slip away
by: Li, Taozhe, et al.
Published: (2026)
by: Li, Taozhe, et al.
Published: (2026)
Don't Waste Bits! Adaptive KV-Cache Quantization for Lightweight On-Device LLMs
by: Boroujeni, Sayed Pedram Haeri, et al.
Published: (2026)
by: Boroujeni, Sayed Pedram Haeri, et al.
Published: (2026)
Don't Judge Before You CLIP: A Unified Approach for Perceptual Tasks
by: Zalcher, Amit, et al.
Published: (2025)
by: Zalcher, Amit, et al.
Published: (2025)
Don't Look at the Camera: Achieving Perceived Eye Contact
by: Gao, Alice, et al.
Published: (2024)
by: Gao, Alice, et al.
Published: (2024)
Don't Fear Peculiar Activation Functions: EUAF and Beyond
by: Wang, Qianchao, et al.
Published: (2024)
by: Wang, Qianchao, et al.
Published: (2024)
Don't Mind the Gaps: Implicit Neural Representations for Resolution-Agnostic Retinal OCT Analysis
by: Kahrs, Bennet, et al.
Published: (2026)
by: Kahrs, Bennet, et al.
Published: (2026)
Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
by: Zou, Xin, et al.
Published: (2025)
by: Zou, Xin, et al.
Published: (2025)
Don't Guess, Just Ask: Resolving Ambiguity in Referring Segmentation via Multi-turn Clarification
by: Yang, Yuting, et al.
Published: (2026)
by: Yang, Yuting, et al.
Published: (2026)
Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models
by: Baid, Ami, et al.
Published: (2026)
by: Baid, Ami, et al.
Published: (2026)
Don't Reach for the Stars: Rethinking Topology for Resilient Federated Learning
by: Konstantin, Mirko, et al.
Published: (2025)
by: Konstantin, Mirko, et al.
Published: (2025)
Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding
by: Kwon, Mincheol, et al.
Published: (2026)
by: Kwon, Mincheol, et al.
Published: (2026)
If At First You Don't Succeed: Test Time Re-ranking for Zero-shot, Cross-domain Retrieval
by: Hudson, Finlay G. C., et al.
Published: (2023)
by: Hudson, Finlay G. C., et al.
Published: (2023)
Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time
by: Cheng, Jintao, et al.
Published: (2025)
by: Cheng, Jintao, et al.
Published: (2025)
Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs
by: Jiao, Pengkun, et al.
Published: (2025)
by: Jiao, Pengkun, et al.
Published: (2025)
Surely Large Multimodal Models (Don't) Excel in Visual Species Recognition?
by: Liu, Tian, et al.
Published: (2025)
by: Liu, Tian, et al.
Published: (2025)
Don't Look Twice: Faster Video Transformers with Run-Length Tokenization
by: Choudhury, Rohan, et al.
Published: (2024)
by: Choudhury, Rohan, et al.
Published: (2024)
Pixels Don't Lie (But Your Detector Might): Bootstrapping MLLM-as-a-Judge for Trustworthy Deepfake Detection and Reasoning Supervision
by: Kuckreja, Kartik, et al.
Published: (2026)
by: Kuckreja, Kartik, et al.
Published: (2026)
Don't Let Your Robot be Harmful: Responsible Robotic Manipulation via Safety-as-Policy
by: Ni, Minheng, et al.
Published: (2024)
by: Ni, Minheng, et al.
Published: (2024)
Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models
by: Li, Senmao, et al.
Published: (2024)
by: Li, Senmao, et al.
Published: (2024)
Now You See It, Now You Don't - Instant Concept Erasure for Safe Text-to-Image and Video Generation
by: Biswas, Shristi Das, et al.
Published: (2025)
by: Biswas, Shristi Das, et al.
Published: (2025)
"Don't forget to put the milk back!" Dataset for Enabling Embodied Agents to Detect Anomalous Situations
by: Mullen Jr, James F., et al.
Published: (2024)
by: Mullen Jr, James F., et al.
Published: (2024)
Render, Don't Decode: Weight-Space World Models with Latent Structural Disentanglement
by: Nzoyem, Roussel Desmond, et al.
Published: (2026)
by: Nzoyem, Roussel Desmond, et al.
Published: (2026)
Don't drop your samples! Coherence-aware training benefits Conditional diffusion
by: Dufour, Nicolas, et al.
Published: (2024)
by: Dufour, Nicolas, et al.
Published: (2024)
Similar Items
-
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
by: Pang, Zhanzhong, et al.
Published: (2026) -
On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action Understanding
by: Pang, Zhanzhong, et al.
Published: (2026) -
Context-Enhanced Memory-Refined Transformer for Online Action Detection
by: Pang, Zhanzhong, et al.
Published: (2025) -
Cost-Sensitive Learning for Long-Tailed Temporal Action Segmentation
by: Pang, Zhanzhong, et al.
Published: (2025) -
Long-Tail Temporal Action Segmentation with Group-wise Temporal Logit Adjustment
by: Pang, Zhanzhong, et al.
Published: (2024)