VidLA: Video-Language Alignment at Scale
Fuente:
arXiv
Salvato in:
| Autori principali: | Rizve, Mamshad Nayeem, Fei, Fan, Unnikrishnan, Jayakrishnan, Tran, Son, Yao, Benjamin Z., Zeng, Belinda, Shah, Mubarak, Chilimbi, Trishul |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Open Vocabulary Multi-Label Video Classification
di: Gupta, Rohit, et al.
Pubblicazione: (2024)
di: Gupta, Rohit, et al.
Pubblicazione: (2024)
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
di: Swetha, Sirnam, et al.
Pubblicazione: (2024)
di: Swetha, Sirnam, et al.
Pubblicazione: (2024)
GAReT: Cross-view Video Geolocalization with Adapters and Auto-Regressive Transformers
di: Pillai, Manu S, et al.
Pubblicazione: (2024)
di: Pillai, Manu S, et al.
Pubblicazione: (2024)
FinePseudo: Improving Pseudo-Labelling through Temporal-Alignablity for Semi-Supervised Fine-Grained Action Recognition
di: Dave, Ishan Rajendrakumar, et al.
Pubblicazione: (2024)
di: Dave, Ishan Rajendrakumar, et al.
Pubblicazione: (2024)
ViLL-E: Video LLM Embeddings for Retrieval
di: Gupta, Rohit, et al.
Pubblicazione: (2026)
di: Gupta, Rohit, et al.
Pubblicazione: (2026)
CompLLM: Compression for Long Context Q&A
di: Berton, Gabriele, et al.
Pubblicazione: (2025)
di: Berton, Gabriele, et al.
Pubblicazione: (2025)
CoLLM: A Large Language Model for Composed Image Retrieval
di: Huynh, Chuong, et al.
Pubblicazione: (2025)
di: Huynh, Chuong, et al.
Pubblicazione: (2025)
Is ImageNet worth 1 video? Learning strong image encoders from 1 long unlabelled video
di: Venkataramanan, Shashanka, et al.
Pubblicazione: (2023)
di: Venkataramanan, Shashanka, et al.
Pubblicazione: (2023)
M-LLM Based Video Frame Selection for Efficient Video Understanding
di: Hu, Kai, et al.
Pubblicazione: (2025)
di: Hu, Kai, et al.
Pubblicazione: (2025)
Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual Generation
di: Kim, Subin, et al.
Pubblicazione: (2025)
di: Kim, Subin, et al.
Pubblicazione: (2025)
CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation
di: Zhu, Zixin, et al.
Pubblicazione: (2025)
di: Zhu, Zixin, et al.
Pubblicazione: (2025)
Unified Alignment Protocol: Making Sense of the Unlabeled Data in New Domains
di: Ahmed, Sabbir, et al.
Pubblicazione: (2025)
di: Ahmed, Sabbir, et al.
Pubblicazione: (2025)
DreamBlend: Advancing Personalized Fine-tuning of Text-to-Image Diffusion Models
di: Ram, Shwetha, et al.
Pubblicazione: (2024)
di: Ram, Shwetha, et al.
Pubblicazione: (2024)
VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale
di: Kulkarni, Parth Parag, et al.
Pubblicazione: (2026)
di: Kulkarni, Parth Parag, et al.
Pubblicazione: (2026)
Evolutionary Contrastive Distillation for Language Model Alignment
di: Katz-Samuels, Julian, et al.
Pubblicazione: (2024)
di: Katz-Samuels, Julian, et al.
Pubblicazione: (2024)
VeRVE: Versatile Retrieval for Videos via Unified Embeddings
di: Halbe, Shaunak, et al.
Pubblicazione: (2026)
di: Halbe, Shaunak, et al.
Pubblicazione: (2026)
Exploring Reasoning-Infused Text Embedding with Large Language Models for Zero-Shot Dense Retrieval
di: Liu, Yuxiang, et al.
Pubblicazione: (2025)
di: Liu, Yuxiang, et al.
Pubblicazione: (2025)
CityGuessr: City-Level Video Geo-Localization on a Global Scale
di: Kulkarni, Parth Parag, et al.
Pubblicazione: (2024)
di: Kulkarni, Parth Parag, et al.
Pubblicazione: (2024)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
di: Qin, Bosheng, et al.
Pubblicazione: (2023)
di: Qin, Bosheng, et al.
Pubblicazione: (2023)
Diffusion Models For Multi-Modal Generative Modeling
di: Chen, Changyou, et al.
Pubblicazione: (2024)
di: Chen, Changyou, et al.
Pubblicazione: (2024)
VIDEOP2R: Video Understanding from Perception to Reasoning
di: Jiang, Yifan, et al.
Pubblicazione: (2025)
di: Jiang, Yifan, et al.
Pubblicazione: (2025)
Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs
di: Luo, Jinqi, et al.
Pubblicazione: (2026)
di: Luo, Jinqi, et al.
Pubblicazione: (2026)
AdaVid: Adaptive Video-Language Pretraining
di: Patel, Chaitanya, et al.
Pubblicazione: (2025)
di: Patel, Chaitanya, et al.
Pubblicazione: (2025)
TimeLogic: A Temporal Logic Benchmark for Video QA
di: Swetha, Sirnam, et al.
Pubblicazione: (2025)
di: Swetha, Sirnam, et al.
Pubblicazione: (2025)
HieraVid: Hierarchical Token Pruning for Fast Video Large Language Models
di: Guo, Yansong, et al.
Pubblicazione: (2026)
di: Guo, Yansong, et al.
Pubblicazione: (2026)
PackCache: A Training-Free Acceleration Method for Unified Autoregressive Video Generation via Compact KV-Cache
di: Li, Kunyang, et al.
Pubblicazione: (2026)
di: Li, Kunyang, et al.
Pubblicazione: (2026)
CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
di: Poppi, Tobia, et al.
Pubblicazione: (2026)
di: Poppi, Tobia, et al.
Pubblicazione: (2026)
Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding
di: Fioresi, Joseph, et al.
Pubblicazione: (2025)
di: Fioresi, Joseph, et al.
Pubblicazione: (2025)
Sync from the Sea: Retrieving Alignable Videos from Large-Scale Datasets
di: Dave, Ishan Rajendrakumar, et al.
Pubblicazione: (2024)
di: Dave, Ishan Rajendrakumar, et al.
Pubblicazione: (2024)
Attend Locally, Remember Linearly: Linear Attention as Cross-Frame Memory for Autoregressive Video Diffusion
di: Li, Kunyang, et al.
Pubblicazione: (2026)
di: Li, Kunyang, et al.
Pubblicazione: (2026)
From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos
di: Gupta, Animesh, et al.
Pubblicazione: (2025)
di: Gupta, Animesh, et al.
Pubblicazione: (2025)
VidText: Towards Comprehensive Evaluation for Video Text Understanding
di: Yang, Zhoufaran, et al.
Pubblicazione: (2025)
di: Yang, Zhoufaran, et al.
Pubblicazione: (2025)
KidLM: Advancing Language Models for Children -- Early Insights and Future Directions
di: Nayeem, Mir Tafseer, et al.
Pubblicazione: (2024)
di: Nayeem, Mir Tafseer, et al.
Pubblicazione: (2024)
DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation
di: Di, Donglin, et al.
Pubblicazione: (2024)
di: Di, Donglin, et al.
Pubblicazione: (2024)
VidCoM: Fast Video Comprehension through Large Language Models with Multimodal Tools
di: Qi, Ji, et al.
Pubblicazione: (2023)
di: Qi, Ji, et al.
Pubblicazione: (2023)
EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation
di: Wang, Xiaofeng, et al.
Pubblicazione: (2024)
di: Wang, Xiaofeng, et al.
Pubblicazione: (2024)
Vid2Coach: Transforming How-To Videos into Task Assistants
di: Huh, Mina, et al.
Pubblicazione: (2025)
di: Huh, Mina, et al.
Pubblicazione: (2025)
HarmoVid: Relightful Video Portrait Harmonization
di: Choi, Jun Myeong, et al.
Pubblicazione: (2026)
di: Choi, Jun Myeong, et al.
Pubblicazione: (2026)
PruneVid: Visual Token Pruning for Efficient Video Large Language Models
di: Huang, Xiaohu, et al.
Pubblicazione: (2024)
di: Huang, Xiaohu, et al.
Pubblicazione: (2024)
EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models
di: Huang, Shiqi, et al.
Pubblicazione: (2026)
di: Huang, Shiqi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Open Vocabulary Multi-Label Video Classification
di: Gupta, Rohit, et al.
Pubblicazione: (2024) -
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
di: Swetha, Sirnam, et al.
Pubblicazione: (2024) -
GAReT: Cross-view Video Geolocalization with Adapters and Auto-Regressive Transformers
di: Pillai, Manu S, et al.
Pubblicazione: (2024) -
FinePseudo: Improving Pseudo-Labelling through Temporal-Alignablity for Semi-Supervised Fine-Grained Action Recognition
di: Dave, Ishan Rajendrakumar, et al.
Pubblicazione: (2024) -
ViLL-E: Video LLM Embeddings for Retrieval
di: Gupta, Rohit, et al.
Pubblicazione: (2026)