Contrastive Language Video Time Pre-training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Hengyue, Min, Kyle, Valdez, Hector A., Tripathi, Subarna |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SViTT-Ego: A Sparse Video-Text Transformer for Egocentric Video
von: Valdez, Hector A., et al.
Veröffentlicht: (2024)
von: Valdez, Hector A., et al.
Veröffentlicht: (2024)
Search2Motion: Training-Free Object-Level Motion Control via Attention-Consensus Search
von: Liu, Sainan, et al.
Veröffentlicht: (2026)
von: Liu, Sainan, et al.
Veröffentlicht: (2026)
Ego-VPA: Egocentric Video Understanding with Parameter-efficient Adaptation
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2024)
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2024)
Harnessing Object Grounding for Time-Sensitive Video Understanding
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2025)
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2025)
Keystep Recognition using Graph Neural Networks
von: Romero, Julia Lee, et al.
Veröffentlicht: (2025)
von: Romero, Julia Lee, et al.
Veröffentlicht: (2025)
Graph-Based Multimodal and Multi-view Alignment for Keystep Recognition
von: Romero, Julia Lee, et al.
Veröffentlicht: (2025)
von: Romero, Julia Lee, et al.
Veröffentlicht: (2025)
VideoSAGE: Video Summarization with Graph Representation Learning
von: Chaves, Jose M. Rojas, et al.
Veröffentlicht: (2024)
von: Chaves, Jose M. Rojas, et al.
Veröffentlicht: (2024)
EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs
von: Rodin, Ivan, et al.
Veröffentlicht: (2025)
von: Rodin, Ivan, et al.
Veröffentlicht: (2025)
VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis
von: Dipta, Shubhashis Roy, et al.
Veröffentlicht: (2025)
von: Dipta, Shubhashis Roy, et al.
Veröffentlicht: (2025)
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2025)
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2025)
TrajPred: Trajectory-Conditioned Joint Embedding Prediction for Surgical Instrument-Tissue Interaction Recognition in Vision-Language Models
von: Cheng, Jiajun, et al.
Veröffentlicht: (2026)
von: Cheng, Jiajun, et al.
Veröffentlicht: (2026)
Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training
von: Wang, Haicheng, et al.
Veröffentlicht: (2024)
von: Wang, Haicheng, et al.
Veröffentlicht: (2024)
PALADIN : Robust Neural Fingerprinting for Text-to-Image Diffusion Models
von: L, Murthy, et al.
Veröffentlicht: (2025)
von: L, Murthy, et al.
Veröffentlicht: (2025)
Comment-aided Video-Language Alignment via Contrastive Pre-training for Short-form Video Humor Detection
von: Liu, Yang, et al.
Veröffentlicht: (2024)
von: Liu, Yang, et al.
Veröffentlicht: (2024)
A Closer Look at the Explainability of Contrastive Language-Image Pre-training
von: Li, Yi, et al.
Veröffentlicht: (2023)
von: Li, Yi, et al.
Veröffentlicht: (2023)
SCAN: Bootstrapping Contrastive Pre-training for Data Efficiency
von: Guo, Yangyang, et al.
Veröffentlicht: (2024)
von: Guo, Yangyang, et al.
Veröffentlicht: (2024)
A Simple Recipe for Contrastively Pre-training Video-First Encoders Beyond 16 Frames
von: Papalampidi, Pinelopi, et al.
Veröffentlicht: (2023)
von: Papalampidi, Pinelopi, et al.
Veröffentlicht: (2023)
Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training
von: Gao, Yipeng, et al.
Veröffentlicht: (2023)
von: Gao, Yipeng, et al.
Veröffentlicht: (2023)
Temporal-Consistent Video Restoration with Pre-trained Diffusion Models
von: Wang, Hengkang, et al.
Veröffentlicht: (2025)
von: Wang, Hengkang, et al.
Veröffentlicht: (2025)
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
von: Jin, Yang, et al.
Veröffentlicht: (2024)
von: Jin, Yang, et al.
Veröffentlicht: (2024)
A Generative Adversarial Approach to Adversarial Attacks Guided by Contrastive Language-Image Pre-trained Model
von: Soor, Sampriti, et al.
Veröffentlicht: (2025)
von: Soor, Sampriti, et al.
Veröffentlicht: (2025)
SCALE-VLP: Soft-Weighted Contrastive Volumetric Vision-Language Pre-training with Spatial-Knowledge Semantics
von: Mahdizadeh, Ailar, et al.
Veröffentlicht: (2025)
von: Mahdizadeh, Ailar, et al.
Veröffentlicht: (2025)
A Multimodal Pre-trained Network for Integrated EEG-Video Seizure Detection
von: Lu, Tong, et al.
Veröffentlicht: (2026)
von: Lu, Tong, et al.
Veröffentlicht: (2026)
Repurposing Pre-trained Video Diffusion Models for Event-based Video Interpolation
von: Chen, Jingxi, et al.
Veröffentlicht: (2024)
von: Chen, Jingxi, et al.
Veröffentlicht: (2024)
Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understanding
von: Jin, Jiayun, et al.
Veröffentlicht: (2026)
von: Jin, Jiayun, et al.
Veröffentlicht: (2026)
Multi-hop Relational Contrastive Learning: Extending Spatial Contrastive Pre-training Beyond Pairwise Relations
von: Ahmed, Sheikh Tanvir, et al.
Veröffentlicht: (2026)
von: Ahmed, Sheikh Tanvir, et al.
Veröffentlicht: (2026)
ByDeWay: Boost Your multimodal LLM with DEpth prompting in a Training-Free Way
von: Roy, Rajarshi, et al.
Veröffentlicht: (2025)
von: Roy, Rajarshi, et al.
Veröffentlicht: (2025)
PatchContrast: Self-Supervised Pre-training for 3D Object Detection
von: Shrout, Oren, et al.
Veröffentlicht: (2023)
von: Shrout, Oren, et al.
Veröffentlicht: (2023)
WeakSupCon: Weakly Supervised Contrastive Learning for Encoder Pre-training
von: Zhang, Bodong, et al.
Veröffentlicht: (2025)
von: Zhang, Bodong, et al.
Veröffentlicht: (2025)
RadCLIP: Enhancing Radiologic Image Analysis through Contrastive Language-Image Pre-training
von: Lu, Zhixiu, et al.
Veröffentlicht: (2024)
von: Lu, Zhixiu, et al.
Veröffentlicht: (2024)
PreMix: Label-Efficient Multiple Instance Learning via Non-Contrastive Pre-training and Feature Mixing
von: Wong, Bryan, et al.
Veröffentlicht: (2024)
von: Wong, Bryan, et al.
Veröffentlicht: (2024)
Large-scale Pre-training for Grounded Video Caption Generation
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2025)
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2025)
3D Scene Graph Guided Vision-Language Pre-training
von: Liu, Hao, et al.
Veröffentlicht: (2024)
von: Liu, Hao, et al.
Veröffentlicht: (2024)
VILA: On Pre-training for Visual Language Models
von: Lin, Ji, et al.
Veröffentlicht: (2023)
von: Lin, Ji, et al.
Veröffentlicht: (2023)
Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation
von: Zhu, Zixin, et al.
Veröffentlicht: (2024)
von: Zhu, Zixin, et al.
Veröffentlicht: (2024)
Training-free Video Temporal Grounding using Large-scale Pre-trained Models
von: Zheng, Minghang, et al.
Veröffentlicht: (2024)
von: Zheng, Minghang, et al.
Veröffentlicht: (2024)
EndoMamba: An Efficient Foundation Model for Endoscopic Videos via Hierarchical Pre-training
von: Tian, Qingyao, et al.
Veröffentlicht: (2025)
von: Tian, Qingyao, et al.
Veröffentlicht: (2025)
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection
von: Zhuang, Weijun, et al.
Veröffentlicht: (2025)
von: Zhuang, Weijun, et al.
Veröffentlicht: (2025)
Foundation Model for Endoscopy Video Analysis via Large-scale Self-supervised Pre-train
von: Wang, Zhao, et al.
Veröffentlicht: (2023)
von: Wang, Zhao, et al.
Veröffentlicht: (2023)
D4C: Data-Free Quantization for Contrastive Language-Image Pre-training Models
von: Zhang, Wenlun, et al.
Veröffentlicht: (2025)
von: Zhang, Wenlun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SViTT-Ego: A Sparse Video-Text Transformer for Egocentric Video
von: Valdez, Hector A., et al.
Veröffentlicht: (2024) -
Search2Motion: Training-Free Object-Level Motion Control via Attention-Consensus Search
von: Liu, Sainan, et al.
Veröffentlicht: (2026) -
Ego-VPA: Egocentric Video Understanding with Parameter-efficient Adaptation
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2024) -
Harnessing Object Grounding for Time-Sensitive Video Understanding
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2025) -
Keystep Recognition using Graph Neural Networks
von: Romero, Julia Lee, et al.
Veröffentlicht: (2025)