ViLL-E: Video LLM Embeddings for Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gupta, Rohit, Unnikrishnan, Jayakrishnan, Fei, Fan, Liu, Sheng, Tran, Son, Shah, Mubarak |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Open Vocabulary Multi-Label Video Classification
von: Gupta, Rohit, et al.
Veröffentlicht: (2024)
von: Gupta, Rohit, et al.
Veröffentlicht: (2024)
VidLA: Video-Language Alignment at Scale
von: Rizve, Mamshad Nayeem, et al.
Veröffentlicht: (2024)
von: Rizve, Mamshad Nayeem, et al.
Veröffentlicht: (2024)
VeRVE: Versatile Retrieval for Videos via Unified Embeddings
von: Halbe, Shaunak, et al.
Veröffentlicht: (2026)
von: Halbe, Shaunak, et al.
Veröffentlicht: (2026)
Cross-View Open-Vocabulary Object Detection in Aerial Imagery
von: Kini, Jyoti, et al.
Veröffentlicht: (2025)
von: Kini, Jyoti, et al.
Veröffentlicht: (2025)
From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos
von: Gupta, Animesh, et al.
Veröffentlicht: (2025)
von: Gupta, Animesh, et al.
Veröffentlicht: (2025)
VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale
von: Kulkarni, Parth Parag, et al.
Veröffentlicht: (2026)
von: Kulkarni, Parth Parag, et al.
Veröffentlicht: (2026)
StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
von: Siddiqui, Nyle, et al.
Veröffentlicht: (2025)
von: Siddiqui, Nyle, et al.
Veröffentlicht: (2025)
CompLLM: Compression for Long Context Q&A
von: Berton, Gabriele, et al.
Veröffentlicht: (2025)
von: Berton, Gabriele, et al.
Veröffentlicht: (2025)
Composed Video Retrieval via Enriched Context and Discriminative Embeddings
von: Thawakar, Omkar, et al.
Veröffentlicht: (2024)
von: Thawakar, Omkar, et al.
Veröffentlicht: (2024)
CoLLM: A Large Language Model for Composed Image Retrieval
von: Huynh, Chuong, et al.
Veröffentlicht: (2025)
von: Huynh, Chuong, et al.
Veröffentlicht: (2025)
BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
von: Narnaware, Vishal, et al.
Veröffentlicht: (2025)
von: Narnaware, Vishal, et al.
Veröffentlicht: (2025)
M-LLM Based Video Frame Selection for Efficient Video Understanding
von: Hu, Kai, et al.
Veröffentlicht: (2025)
von: Hu, Kai, et al.
Veröffentlicht: (2025)
The Telephone Game: Evaluating Semantic Drift in Unified Models
von: Mollah, Sabbir, et al.
Veröffentlicht: (2025)
von: Mollah, Sabbir, et al.
Veröffentlicht: (2025)
TimeLogic: A Temporal Logic Benchmark for Video QA
von: Swetha, Sirnam, et al.
Veröffentlicht: (2025)
von: Swetha, Sirnam, et al.
Veröffentlicht: (2025)
Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos
von: Gupta, Rohit, et al.
Veröffentlicht: (2025)
von: Gupta, Rohit, et al.
Veröffentlicht: (2025)
Investigating Memorization in Video Diffusion Models
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
von: Swetha, Sirnam, et al.
Veröffentlicht: (2025)
von: Swetha, Sirnam, et al.
Veröffentlicht: (2025)
PackCache: A Training-Free Acceleration Method for Unified Autoregressive Video Generation via Compact KV-Cache
von: Li, Kunyang, et al.
Veröffentlicht: (2026)
von: Li, Kunyang, et al.
Veröffentlicht: (2026)
Beyond Simple Edits: Composed Video Retrieval with Dense Modifications
von: Thawakar, Omkar, et al.
Veröffentlicht: (2025)
von: Thawakar, Omkar, et al.
Veröffentlicht: (2025)
Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding
von: Fioresi, Joseph, et al.
Veröffentlicht: (2025)
von: Fioresi, Joseph, et al.
Veröffentlicht: (2025)
GAReT: Cross-view Video Geolocalization with Adapters and Auto-Regressive Transformers
von: Pillai, Manu S, et al.
Veröffentlicht: (2024)
von: Pillai, Manu S, et al.
Veröffentlicht: (2024)
CityGuessr: City-Level Video Geo-Localization on a Global Scale
von: Kulkarni, Parth Parag, et al.
Veröffentlicht: (2024)
von: Kulkarni, Parth Parag, et al.
Veröffentlicht: (2024)
GT-Loc: Unifying When and Where in Images Through a Joint Embedding Space
von: Shatwell, David G., et al.
Veröffentlicht: (2025)
von: Shatwell, David G., et al.
Veröffentlicht: (2025)
VIDEOP2R: Video Understanding from Perception to Reasoning
von: Jiang, Yifan, et al.
Veröffentlicht: (2025)
von: Jiang, Yifan, et al.
Veröffentlicht: (2025)
Leveraging Pre-Trained Visual Models for AI-Generated Video Detection
von: Veeramachaneni, Keerthi, et al.
Veröffentlicht: (2025)
von: Veeramachaneni, Keerthi, et al.
Veröffentlicht: (2025)
LL-ViT: Edge Deployable Vision Transformers with Look Up Table Neurons
von: Nag, Shashank, et al.
Veröffentlicht: (2025)
von: Nag, Shashank, et al.
Veröffentlicht: (2025)
Attend Locally, Remember Linearly: Linear Attention as Cross-Frame Memory for Autoregressive Video Diffusion
von: Li, Kunyang, et al.
Veröffentlicht: (2026)
von: Li, Kunyang, et al.
Veröffentlicht: (2026)
Temporally Consistent Referring Video Object Segmentation with Hybrid Memory
von: Miao, Bo, et al.
Veröffentlicht: (2024)
von: Miao, Bo, et al.
Veröffentlicht: (2024)
Sync from the Sea: Retrieving Alignable Videos from Large-Scale Datasets
von: Dave, Ishan Rajendrakumar, et al.
Veröffentlicht: (2024)
von: Dave, Ishan Rajendrakumar, et al.
Veröffentlicht: (2024)
Seeing to Ground: Visual Attention for Hallucination-Resilient MDLLMs
von: Narnaware, Vishal, et al.
Veröffentlicht: (2026)
von: Narnaware, Vishal, et al.
Veröffentlicht: (2026)
TIGeR: A Unified Framework for Time, Images and Geo-location Retrieval
von: Shatwell, David G., et al.
Veröffentlicht: (2026)
von: Shatwell, David G., et al.
Veröffentlicht: (2026)
ViViD: Video Virtual Try-on using Diffusion Models
von: Fang, Zixun, et al.
Veröffentlicht: (2024)
von: Fang, Zixun, et al.
Veröffentlicht: (2024)
Learnability-Guided Diffusion for Dataset Distillation
von: Chan-Santiago, Jeffrey A., et al.
Veröffentlicht: (2026)
von: Chan-Santiago, Jeffrey A., et al.
Veröffentlicht: (2026)
MobileViCLIP: An Efficient Video-Text Model for Mobile Devices
von: Yang, Min, et al.
Veröffentlicht: (2025)
von: Yang, Min, et al.
Veröffentlicht: (2025)
MoViE: Mobile Diffusion for Video Editing
von: Karjauv, Adil, et al.
Veröffentlicht: (2024)
von: Karjauv, Adil, et al.
Veröffentlicht: (2024)
Enhancing Privacy-Utility Trade-offs to Mitigate Memorization in Diffusion Models
von: Chen, Chen, et al.
Veröffentlicht: (2025)
von: Chen, Chen, et al.
Veröffentlicht: (2025)
SegVG: Transferring Object Bounding Box to Segmentation for Visual Grounding
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
Exploring Local Memorization in Diffusion Models via Bright Ending Attention
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval
von: Wang, Jiamian, et al.
Veröffentlicht: (2024)
von: Wang, Jiamian, et al.
Veröffentlicht: (2024)
GVD: Guiding Video Diffusion Model for Scalable Video Distillation
von: Li, Kunyang, et al.
Veröffentlicht: (2025)
von: Li, Kunyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Open Vocabulary Multi-Label Video Classification
von: Gupta, Rohit, et al.
Veröffentlicht: (2024) -
VidLA: Video-Language Alignment at Scale
von: Rizve, Mamshad Nayeem, et al.
Veröffentlicht: (2024) -
VeRVE: Versatile Retrieval for Videos via Unified Embeddings
von: Halbe, Shaunak, et al.
Veröffentlicht: (2026) -
Cross-View Open-Vocabulary Object Detection in Aerial Imagery
von: Kini, Jyoti, et al.
Veröffentlicht: (2025) -
From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos
von: Gupta, Animesh, et al.
Veröffentlicht: (2025)