CoRDS: Coreset-based Representative and Diverse Selection for Streaming Video Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Mahdizadeh, Ailar, Azadi, Puria, Li, Muchen, He, Xiangteng, Sigal, Leonid |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SCALE-VLP: Soft-Weighted Contrastive Volumetric Vision-Language Pre-training with Spatial-Knowledge Semantics
di: Mahdizadeh, Ailar, et al.
Pubblicazione: (2025)
di: Mahdizadeh, Ailar, et al.
Pubblicazione: (2025)
InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models
di: Sakai, Shunsuke, et al.
Pubblicazione: (2025)
di: Sakai, Shunsuke, et al.
Pubblicazione: (2025)
DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
di: He, Xiangteng, et al.
Pubblicazione: (2025)
di: He, Xiangteng, et al.
Pubblicazione: (2025)
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
di: Rahman, Tanzila, et al.
Pubblicazione: (2026)
di: Rahman, Tanzila, et al.
Pubblicazione: (2026)
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
di: Luo, Jiayun, et al.
Pubblicazione: (2025)
di: Luo, Jiayun, et al.
Pubblicazione: (2025)
Advancing Complex Wide-Area Scene Understanding with Hierarchical Coresets Selection
di: Wang, Jingyao, et al.
Pubblicazione: (2025)
di: Wang, Jingyao, et al.
Pubblicazione: (2025)
Factorized Video Autoencoders for Efficient Generative Modelling
di: Suhail, Mohammed, et al.
Pubblicazione: (2024)
di: Suhail, Mohammed, et al.
Pubblicazione: (2024)
Implicit and Explicit Commonsense for Multi-sentence Video Captioning
di: Chou, Shih-Han, et al.
Pubblicazione: (2023)
di: Chou, Shih-Han, et al.
Pubblicazione: (2023)
StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
di: Wang, Junxi, et al.
Pubblicazione: (2026)
di: Wang, Junxi, et al.
Pubblicazione: (2026)
A Coreset Selection of Coreset Selection Literature: Introduction and Recent Advances
di: Moser, Brian B., et al.
Pubblicazione: (2025)
di: Moser, Brian B., et al.
Pubblicazione: (2025)
Representing Animatable Avatar via Factorized Neural Fields
di: Song, Chunjin, et al.
Pubblicazione: (2024)
di: Song, Chunjin, et al.
Pubblicazione: (2024)
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
di: Goyal, Raghav, et al.
Pubblicazione: (2023)
di: Goyal, Raghav, et al.
Pubblicazione: (2023)
Efficient and Effective In-context Demonstration Selection with Coreset
di: Wang, Zihua, et al.
Pubblicazione: (2025)
di: Wang, Zihua, et al.
Pubblicazione: (2025)
Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs
di: Xian, Jia Jun Cheng, et al.
Pubblicazione: (2025)
di: Xian, Jia Jun Cheng, et al.
Pubblicazione: (2025)
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
di: Chinchure, Aditya, et al.
Pubblicazione: (2025)
di: Chinchure, Aditya, et al.
Pubblicazione: (2025)
StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding
di: Yang, Haolin, et al.
Pubblicazione: (2025)
di: Yang, Haolin, et al.
Pubblicazione: (2025)
Coreset Selection for Object Detection
di: Lee, Hojun, et al.
Pubblicazione: (2024)
di: Lee, Hojun, et al.
Pubblicazione: (2024)
StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
di: Wang, Yifei, et al.
Pubblicazione: (2025)
di: Wang, Yifei, et al.
Pubblicazione: (2025)
SPIKE-RL: Video-LLMs meet Bayesian Surprise
di: Ravi, Sahithya, et al.
Pubblicazione: (2025)
di: Ravi, Sahithya, et al.
Pubblicazione: (2025)
On the Fairness, Diversity and Reliability of Text-to-Image Generative Models
di: Vice, Jordan, et al.
Pubblicazione: (2024)
di: Vice, Jordan, et al.
Pubblicazione: (2024)
Towards Video Anomaly Retrieval from Video Anomaly Detection: New Benchmarks and Model
di: Wu, Peng, et al.
Pubblicazione: (2023)
di: Wu, Peng, et al.
Pubblicazione: (2023)
Zero-Shot Coreset Selection via Iterative Subspace Sampling
di: Griffin, Brent A., et al.
Pubblicazione: (2024)
di: Griffin, Brent A., et al.
Pubblicazione: (2024)
Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos
di: Pan, Yulin, et al.
Pubblicazione: (2023)
di: Pan, Yulin, et al.
Pubblicazione: (2023)
An Efficient Streaming Video Understanding Framework with Agentic Control
di: Liu, Jinming, et al.
Pubblicazione: (2026)
di: Liu, Jinming, et al.
Pubblicazione: (2026)
Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities
di: Chandhok, Shivam, et al.
Pubblicazione: (2024)
di: Chandhok, Shivam, et al.
Pubblicazione: (2024)
Evolution-aware VAriance (EVA) Coreset Selection for Medical Image Classification
di: Hong, Yuxin, et al.
Pubblicazione: (2024)
di: Hong, Yuxin, et al.
Pubblicazione: (2024)
Multi-modal News Understanding with Professionally Labelled Videos (ReutersViLNews)
di: Chou, Shih-Han, et al.
Pubblicazione: (2024)
di: Chou, Shih-Han, et al.
Pubblicazione: (2024)
DeformStream: Deformation-based Adaptive Volumetric Video Streaming
di: Li, Boyan, et al.
Pubblicazione: (2024)
di: Li, Boyan, et al.
Pubblicazione: (2024)
StreamingVLM: Real-Time Understanding for Infinite Video Streams
di: Xu, Ruyi, et al.
Pubblicazione: (2025)
di: Xu, Ruyi, et al.
Pubblicazione: (2025)
Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events
di: Chinchure, Aditya, et al.
Pubblicazione: (2024)
di: Chinchure, Aditya, et al.
Pubblicazione: (2024)
CAST: Collapse-Aware multi-Scale Topology Fusion for Multimodal Coreset Selection
di: Zhao, Boran, et al.
Pubblicazione: (2026)
di: Zhao, Boran, et al.
Pubblicazione: (2026)
The Impact of Coreset Selection on Spurious Correlations and Group Robustness
di: Dharmasiri, Amaya, et al.
Pubblicazione: (2025)
di: Dharmasiri, Amaya, et al.
Pubblicazione: (2025)
Streaming Long Video Understanding with Large Language Models
di: Qian, Rui, et al.
Pubblicazione: (2024)
di: Qian, Rui, et al.
Pubblicazione: (2024)
VideoScan: Enabling Efficient Streaming Video Understanding via Frame-level Semantic Carriers
di: Li, Ruanjun, et al.
Pubblicazione: (2025)
di: Li, Ruanjun, et al.
Pubblicazione: (2025)
StreamForest: Efficient Online Video Understanding with Persistent Event Memory
di: Zeng, Xiangyu, et al.
Pubblicazione: (2025)
di: Zeng, Xiangyu, et al.
Pubblicazione: (2025)
AURA: Always-On Understanding and Real-Time Assistance via Video Streams
di: Lu, Xudong, et al.
Pubblicazione: (2026)
di: Lu, Xudong, et al.
Pubblicazione: (2026)
Preventing Catastrophic Forgetting through Memory Networks in Continuous Detection
di: Bhatt, Gaurav, et al.
Pubblicazione: (2024)
di: Bhatt, Gaurav, et al.
Pubblicazione: (2024)
Enhancing Semi-Supervised Learning via Representative and Diverse Sample Selection
di: Shao, Qian, et al.
Pubblicazione: (2024)
di: Shao, Qian, et al.
Pubblicazione: (2024)
CurveStream: Boosting Streaming Video Understanding in MLLMs via Curvature-Aware Hierarchical Visual Memory Management
di: Wang, Chao, et al.
Pubblicazione: (2026)
di: Wang, Chao, et al.
Pubblicazione: (2026)
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
di: Chatterjee, Dibyadip, et al.
Pubblicazione: (2025)
di: Chatterjee, Dibyadip, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SCALE-VLP: Soft-Weighted Contrastive Volumetric Vision-Language Pre-training with Spatial-Knowledge Semantics
di: Mahdizadeh, Ailar, et al.
Pubblicazione: (2025) -
InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models
di: Sakai, Shunsuke, et al.
Pubblicazione: (2025) -
DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
di: He, Xiangteng, et al.
Pubblicazione: (2025) -
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
di: Rahman, Tanzila, et al.
Pubblicazione: (2026) -
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
di: Luo, Jiayun, et al.
Pubblicazione: (2025)