Open Vocabulary Multi-Label Video Classification
Fuente:
arXiv
Salvato in:
| Autori principali: | Gupta, Rohit, Rizve, Mamshad Nayeem, Unnikrishnan, Jayakrishnan, Tawari, Ashish, Tran, Son, Shah, Mubarak, Yao, Benjamin, Chilimbi, Trishul |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VidLA: Video-Language Alignment at Scale
di: Rizve, Mamshad Nayeem, et al.
Pubblicazione: (2024)
di: Rizve, Mamshad Nayeem, et al.
Pubblicazione: (2024)
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
di: Swetha, Sirnam, et al.
Pubblicazione: (2024)
di: Swetha, Sirnam, et al.
Pubblicazione: (2024)
GAReT: Cross-view Video Geolocalization with Adapters and Auto-Regressive Transformers
di: Pillai, Manu S, et al.
Pubblicazione: (2024)
di: Pillai, Manu S, et al.
Pubblicazione: (2024)
FinePseudo: Improving Pseudo-Labelling through Temporal-Alignablity for Semi-Supervised Fine-Grained Action Recognition
di: Dave, Ishan Rajendrakumar, et al.
Pubblicazione: (2024)
di: Dave, Ishan Rajendrakumar, et al.
Pubblicazione: (2024)
CoLLM: A Large Language Model for Composed Image Retrieval
di: Huynh, Chuong, et al.
Pubblicazione: (2025)
di: Huynh, Chuong, et al.
Pubblicazione: (2025)
ViLL-E: Video LLM Embeddings for Retrieval
di: Gupta, Rohit, et al.
Pubblicazione: (2026)
di: Gupta, Rohit, et al.
Pubblicazione: (2026)
Cross-View Open-Vocabulary Object Detection in Aerial Imagery
di: Kini, Jyoti, et al.
Pubblicazione: (2025)
di: Kini, Jyoti, et al.
Pubblicazione: (2025)
Is ImageNet worth 1 video? Learning strong image encoders from 1 long unlabelled video
di: Venkataramanan, Shashanka, et al.
Pubblicazione: (2023)
di: Venkataramanan, Shashanka, et al.
Pubblicazione: (2023)
M-LLM Based Video Frame Selection for Efficient Video Understanding
di: Hu, Kai, et al.
Pubblicazione: (2025)
di: Hu, Kai, et al.
Pubblicazione: (2025)
CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation
di: Zhu, Zixin, et al.
Pubblicazione: (2025)
di: Zhu, Zixin, et al.
Pubblicazione: (2025)
DreamBlend: Advancing Personalized Fine-tuning of Text-to-Image Diffusion Models
di: Ram, Shwetha, et al.
Pubblicazione: (2024)
di: Ram, Shwetha, et al.
Pubblicazione: (2024)
Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual Generation
di: Kim, Subin, et al.
Pubblicazione: (2025)
di: Kim, Subin, et al.
Pubblicazione: (2025)
CompLLM: Compression for Long Context Q&A
di: Berton, Gabriele, et al.
Pubblicazione: (2025)
di: Berton, Gabriele, et al.
Pubblicazione: (2025)
Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos
di: Gupta, Rohit, et al.
Pubblicazione: (2025)
di: Gupta, Rohit, et al.
Pubblicazione: (2025)
Unified Alignment Protocol: Making Sense of the Unlabeled Data in New Domains
di: Ahmed, Sabbir, et al.
Pubblicazione: (2025)
di: Ahmed, Sabbir, et al.
Pubblicazione: (2025)
VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale
di: Kulkarni, Parth Parag, et al.
Pubblicazione: (2026)
di: Kulkarni, Parth Parag, et al.
Pubblicazione: (2026)
Query-Based Knowledge Sharing for Open-Vocabulary Multi-Label Classification
di: Zhu, Xuelin, et al.
Pubblicazione: (2024)
di: Zhu, Xuelin, et al.
Pubblicazione: (2024)
StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
di: Siddiqui, Nyle, et al.
Pubblicazione: (2025)
di: Siddiqui, Nyle, et al.
Pubblicazione: (2025)
VeRVE: Versatile Retrieval for Videos via Unified Embeddings
di: Halbe, Shaunak, et al.
Pubblicazione: (2026)
di: Halbe, Shaunak, et al.
Pubblicazione: (2026)
From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos
di: Gupta, Animesh, et al.
Pubblicazione: (2025)
di: Gupta, Animesh, et al.
Pubblicazione: (2025)
BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
di: Narnaware, Vishal, et al.
Pubblicazione: (2025)
di: Narnaware, Vishal, et al.
Pubblicazione: (2025)
DART: Dual Adaptive Refinement Transfer for Open-Vocabulary Multi-Label Recognition
di: Liu, Haijing, et al.
Pubblicazione: (2025)
di: Liu, Haijing, et al.
Pubblicazione: (2025)
TimeLogic: A Temporal Logic Benchmark for Video QA
di: Swetha, Sirnam, et al.
Pubblicazione: (2025)
di: Swetha, Sirnam, et al.
Pubblicazione: (2025)
Multi-output Deep-Supervised Classifier Chains for Plant Pathology
di: Yao, Jianping, et al.
Pubblicazione: (2025)
di: Yao, Jianping, et al.
Pubblicazione: (2025)
VIDEOP2R: Video Understanding from Perception to Reasoning
di: Jiang, Yifan, et al.
Pubblicazione: (2025)
di: Jiang, Yifan, et al.
Pubblicazione: (2025)
Open-Vocabulary Video Anomaly Detection
di: Wu, Peng, et al.
Pubblicazione: (2023)
di: Wu, Peng, et al.
Pubblicazione: (2023)
The Telephone Game: Evaluating Semantic Drift in Unified Models
di: Mollah, Sabbir, et al.
Pubblicazione: (2025)
di: Mollah, Sabbir, et al.
Pubblicazione: (2025)
VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
di: Swetha, Sirnam, et al.
Pubblicazione: (2025)
di: Swetha, Sirnam, et al.
Pubblicazione: (2025)
PackCache: A Training-Free Acceleration Method for Unified Autoregressive Video Generation via Compact KV-Cache
di: Li, Kunyang, et al.
Pubblicazione: (2026)
di: Li, Kunyang, et al.
Pubblicazione: (2026)
Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding
di: Fioresi, Joseph, et al.
Pubblicazione: (2025)
di: Fioresi, Joseph, et al.
Pubblicazione: (2025)
Towards Open-Vocabulary Semantic Segmentation Without Semantic Labels
di: Shin, Heeseong, et al.
Pubblicazione: (2024)
di: Shin, Heeseong, et al.
Pubblicazione: (2024)
Recover and Match: Open-Vocabulary Multi-Label Recognition through Knowledge-Constrained Optimal Transport
di: Tan, Hao, et al.
Pubblicazione: (2025)
di: Tan, Hao, et al.
Pubblicazione: (2025)
Category-Adaptive Cross-Modal Semantic Refinement and Transfer for Open-Vocabulary Multi-Label Recognition
di: Liu, Haijing, et al.
Pubblicazione: (2024)
di: Liu, Haijing, et al.
Pubblicazione: (2024)
Unsupervised Open-Vocabulary Object Localization in Videos
di: Fan, Ke, et al.
Pubblicazione: (2023)
di: Fan, Ke, et al.
Pubblicazione: (2023)
FOLK: Fast Open-Vocabulary 3D Instance Segmentation via Label-guided Knowledge Distillation
di: Wu, Hongrui, et al.
Pubblicazione: (2025)
di: Wu, Hongrui, et al.
Pubblicazione: (2025)
Attend Locally, Remember Linearly: Linear Attention as Cross-Frame Memory for Autoregressive Video Diffusion
di: Li, Kunyang, et al.
Pubblicazione: (2026)
di: Li, Kunyang, et al.
Pubblicazione: (2026)
Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models
di: Rahman, Muhammad Atta ur, et al.
Pubblicazione: (2025)
di: Rahman, Muhammad Atta ur, et al.
Pubblicazione: (2025)
CityGuessr: City-Level Video Geo-Localization on a Global Scale
di: Kulkarni, Parth Parag, et al.
Pubblicazione: (2024)
di: Kulkarni, Parth Parag, et al.
Pubblicazione: (2024)
Anomize: Better Open Vocabulary Video Anomaly Detection
di: Li, Fei, et al.
Pubblicazione: (2025)
di: Li, Fei, et al.
Pubblicazione: (2025)
Anytime Continual Learning for Open Vocabulary Classification
di: Zhu, Zhen, et al.
Pubblicazione: (2024)
di: Zhu, Zhen, et al.
Pubblicazione: (2024)
Documenti analoghi
-
VidLA: Video-Language Alignment at Scale
di: Rizve, Mamshad Nayeem, et al.
Pubblicazione: (2024) -
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
di: Swetha, Sirnam, et al.
Pubblicazione: (2024) -
GAReT: Cross-view Video Geolocalization with Adapters and Auto-Regressive Transformers
di: Pillai, Manu S, et al.
Pubblicazione: (2024) -
FinePseudo: Improving Pseudo-Labelling through Temporal-Alignablity for Semi-Supervised Fine-Grained Action Recognition
di: Dave, Ishan Rajendrakumar, et al.
Pubblicazione: (2024) -
CoLLM: A Large Language Model for Composed Image Retrieval
di: Huynh, Chuong, et al.
Pubblicazione: (2025)