CoRDS: Coreset-based Representative and Diverse Selection for Streaming Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Mahdizadeh, Ailar, Azadi, Puria, Li, Muchen, He, Xiangteng, Sigal, Leonid |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SCALE-VLP: Soft-Weighted Contrastive Volumetric Vision-Language Pre-training with Spatial-Knowledge Semantics
by: Mahdizadeh, Ailar, et al.
Published: (2025)
by: Mahdizadeh, Ailar, et al.
Published: (2025)
InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models
by: Sakai, Shunsuke, et al.
Published: (2025)
by: Sakai, Shunsuke, et al.
Published: (2025)
DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
by: He, Xiangteng, et al.
Published: (2025)
by: He, Xiangteng, et al.
Published: (2025)
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
by: Rahman, Tanzila, et al.
Published: (2026)
by: Rahman, Tanzila, et al.
Published: (2026)
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
by: Luo, Jiayun, et al.
Published: (2025)
by: Luo, Jiayun, et al.
Published: (2025)
Advancing Complex Wide-Area Scene Understanding with Hierarchical Coresets Selection
by: Wang, Jingyao, et al.
Published: (2025)
by: Wang, Jingyao, et al.
Published: (2025)
Factorized Video Autoencoders for Efficient Generative Modelling
by: Suhail, Mohammed, et al.
Published: (2024)
by: Suhail, Mohammed, et al.
Published: (2024)
Implicit and Explicit Commonsense for Multi-sentence Video Captioning
by: Chou, Shih-Han, et al.
Published: (2023)
by: Chou, Shih-Han, et al.
Published: (2023)
StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
by: Wang, Junxi, et al.
Published: (2026)
by: Wang, Junxi, et al.
Published: (2026)
A Coreset Selection of Coreset Selection Literature: Introduction and Recent Advances
by: Moser, Brian B., et al.
Published: (2025)
by: Moser, Brian B., et al.
Published: (2025)
Representing Animatable Avatar via Factorized Neural Fields
by: Song, Chunjin, et al.
Published: (2024)
by: Song, Chunjin, et al.
Published: (2024)
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
by: Goyal, Raghav, et al.
Published: (2023)
by: Goyal, Raghav, et al.
Published: (2023)
Efficient and Effective In-context Demonstration Selection with Coreset
by: Wang, Zihua, et al.
Published: (2025)
by: Wang, Zihua, et al.
Published: (2025)
Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs
by: Xian, Jia Jun Cheng, et al.
Published: (2025)
by: Xian, Jia Jun Cheng, et al.
Published: (2025)
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
by: Chinchure, Aditya, et al.
Published: (2025)
by: Chinchure, Aditya, et al.
Published: (2025)
StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
Coreset Selection for Object Detection
by: Lee, Hojun, et al.
Published: (2024)
by: Lee, Hojun, et al.
Published: (2024)
StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
by: Wang, Yifei, et al.
Published: (2025)
by: Wang, Yifei, et al.
Published: (2025)
SPIKE-RL: Video-LLMs meet Bayesian Surprise
by: Ravi, Sahithya, et al.
Published: (2025)
by: Ravi, Sahithya, et al.
Published: (2025)
On the Fairness, Diversity and Reliability of Text-to-Image Generative Models
by: Vice, Jordan, et al.
Published: (2024)
by: Vice, Jordan, et al.
Published: (2024)
Towards Video Anomaly Retrieval from Video Anomaly Detection: New Benchmarks and Model
by: Wu, Peng, et al.
Published: (2023)
by: Wu, Peng, et al.
Published: (2023)
Zero-Shot Coreset Selection via Iterative Subspace Sampling
by: Griffin, Brent A., et al.
Published: (2024)
by: Griffin, Brent A., et al.
Published: (2024)
Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos
by: Pan, Yulin, et al.
Published: (2023)
by: Pan, Yulin, et al.
Published: (2023)
An Efficient Streaming Video Understanding Framework with Agentic Control
by: Liu, Jinming, et al.
Published: (2026)
by: Liu, Jinming, et al.
Published: (2026)
Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities
by: Chandhok, Shivam, et al.
Published: (2024)
by: Chandhok, Shivam, et al.
Published: (2024)
Evolution-aware VAriance (EVA) Coreset Selection for Medical Image Classification
by: Hong, Yuxin, et al.
Published: (2024)
by: Hong, Yuxin, et al.
Published: (2024)
Multi-modal News Understanding with Professionally Labelled Videos (ReutersViLNews)
by: Chou, Shih-Han, et al.
Published: (2024)
by: Chou, Shih-Han, et al.
Published: (2024)
DeformStream: Deformation-based Adaptive Volumetric Video Streaming
by: Li, Boyan, et al.
Published: (2024)
by: Li, Boyan, et al.
Published: (2024)
StreamingVLM: Real-Time Understanding for Infinite Video Streams
by: Xu, Ruyi, et al.
Published: (2025)
by: Xu, Ruyi, et al.
Published: (2025)
Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events
by: Chinchure, Aditya, et al.
Published: (2024)
by: Chinchure, Aditya, et al.
Published: (2024)
CAST: Collapse-Aware multi-Scale Topology Fusion for Multimodal Coreset Selection
by: Zhao, Boran, et al.
Published: (2026)
by: Zhao, Boran, et al.
Published: (2026)
The Impact of Coreset Selection on Spurious Correlations and Group Robustness
by: Dharmasiri, Amaya, et al.
Published: (2025)
by: Dharmasiri, Amaya, et al.
Published: (2025)
Streaming Long Video Understanding with Large Language Models
by: Qian, Rui, et al.
Published: (2024)
by: Qian, Rui, et al.
Published: (2024)
VideoScan: Enabling Efficient Streaming Video Understanding via Frame-level Semantic Carriers
by: Li, Ruanjun, et al.
Published: (2025)
by: Li, Ruanjun, et al.
Published: (2025)
StreamForest: Efficient Online Video Understanding with Persistent Event Memory
by: Zeng, Xiangyu, et al.
Published: (2025)
by: Zeng, Xiangyu, et al.
Published: (2025)
AURA: Always-On Understanding and Real-Time Assistance via Video Streams
by: Lu, Xudong, et al.
Published: (2026)
by: Lu, Xudong, et al.
Published: (2026)
Preventing Catastrophic Forgetting through Memory Networks in Continuous Detection
by: Bhatt, Gaurav, et al.
Published: (2024)
by: Bhatt, Gaurav, et al.
Published: (2024)
Enhancing Semi-Supervised Learning via Representative and Diverse Sample Selection
by: Shao, Qian, et al.
Published: (2024)
by: Shao, Qian, et al.
Published: (2024)
CurveStream: Boosting Streaming Video Understanding in MLLMs via Curvature-Aware Hierarchical Visual Memory Management
by: Wang, Chao, et al.
Published: (2026)
by: Wang, Chao, et al.
Published: (2026)
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
by: Chatterjee, Dibyadip, et al.
Published: (2025)
by: Chatterjee, Dibyadip, et al.
Published: (2025)
Similar Items
-
SCALE-VLP: Soft-Weighted Contrastive Volumetric Vision-Language Pre-training with Spatial-Knowledge Semantics
by: Mahdizadeh, Ailar, et al.
Published: (2025) -
InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models
by: Sakai, Shunsuke, et al.
Published: (2025) -
DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
by: He, Xiangteng, et al.
Published: (2025) -
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
by: Rahman, Tanzila, et al.
Published: (2026) -
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
by: Luo, Jiayun, et al.
Published: (2025)