A Simple Recipe for Contrastively Pre-training Video-First Encoders Beyond 16 Frames
Fuente:
arXiv
Saved in:
| Main Authors: | Papalampidi, Pinelopi, Koppula, Skanda, Pathak, Shreya, Chiu, Justin, Heyward, Joe, Patraucean, Viorica, Shen, Jiajun, Miech, Antoine, Zisserman, Andrew, Nematzadeh, Aida |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
by: Heyward, Joseph, et al.
Published: (2024)
by: Heyward, Joseph, et al.
Published: (2024)
Memory Consolidation Enables Long-Context Video Understanding
by: Balažević, Ivana, et al.
Published: (2024)
by: Balažević, Ivana, et al.
Published: (2024)
TAPVid-3D: A Benchmark for Tracking Any Point in 3D
by: Koppula, Skanda, et al.
Published: (2024)
by: Koppula, Skanda, et al.
Published: (2024)
Perception Test 2025: Challenge Summary and a Unified VQA Extension
by: Heyward, Joseph, et al.
Published: (2026)
by: Heyward, Joseph, et al.
Published: (2026)
Dynamic Classifier-Free Diffusion Guidance via Online Feedback
by: Papalampidi, Pinelopi, et al.
Published: (2025)
by: Papalampidi, Pinelopi, et al.
Published: (2025)
Learning from Streaming Video with Orthogonal Gradients
by: Han, Tengda, et al.
Published: (2025)
by: Han, Tengda, et al.
Published: (2025)
Finding the Right Moment: Human-Assisted Trailer Creation via Task Composition
by: Papalampidi, Pinelopi, et al.
Published: (2021)
by: Papalampidi, Pinelopi, et al.
Published: (2021)
TAPNext: Tracking Any Point (TAP) as Next Token Prediction
by: Zholus, Artem, et al.
Published: (2025)
by: Zholus, Artem, et al.
Published: (2025)
BootsTAP: Bootstrapped Training for Tracking-Any-Point
by: Doersch, Carl, et al.
Published: (2024)
by: Doersch, Carl, et al.
Published: (2024)
Learning from One Continuous Video Stream
by: Carreira, João, et al.
Published: (2023)
by: Carreira, João, et al.
Published: (2023)
A Mixed Diet Makes DINO An Omnivorous Vision Encoder
by: Kabra, Rishabh, et al.
Published: (2026)
by: Kabra, Rishabh, et al.
Published: (2026)
SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications
by: Hasson, Yana, et al.
Published: (2025)
by: Hasson, Yana, et al.
Published: (2025)
How Overconfidence in Initial Choices and Underconfidence Under Criticism Modulate Change of Mind in Large Language Models
by: Kumaran, Dharshan, et al.
Published: (2025)
by: Kumaran, Dharshan, et al.
Published: (2025)
How LLMs Detect and Correct Their Own Errors: The Role of Internal Confidence Signals
by: Kumaran, Dharshan, et al.
Published: (2026)
by: Kumaran, Dharshan, et al.
Published: (2026)
Causal Evidence that Language Models use Confidence to Drive Behavior
by: Kumaran, Dharshan, et al.
Published: (2026)
by: Kumaran, Dharshan, et al.
Published: (2026)
Dynamic Reflections: Probing Video Representations with Text Alignment
by: Zhu, Tyler, et al.
Published: (2025)
by: Zhu, Tyler, et al.
Published: (2025)
Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings
by: Wiles, Olivia, et al.
Published: (2024)
by: Wiles, Olivia, et al.
Published: (2024)
Recipes for Pre-training LLMs with MXFP8
by: Mishra, Asit, et al.
Published: (2025)
by: Mishra, Asit, et al.
Published: (2025)
How do LLMs Compute Verbal Confidence
by: Kumaran, Dharshan, et al.
Published: (2026)
by: Kumaran, Dharshan, et al.
Published: (2026)
WeakSupCon: Weakly Supervised Contrastive Learning for Encoder Pre-training
by: Zhang, Bodong, et al.
Published: (2025)
by: Zhang, Bodong, et al.
Published: (2025)
Unique Lives, Shared World: Learning from Single-Life Videos
by: Han, Tengda, et al.
Published: (2025)
by: Han, Tengda, et al.
Published: (2025)
Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM Overthinking
by: Zhang, Xinliang Frederick, et al.
Published: (2025)
by: Zhang, Xinliang Frederick, et al.
Published: (2025)
Failed schemes of relatedness in domestic work: Filipina domestic workers in Greece
by: Pinelopi Topali
Published: (2024)
by: Pinelopi Topali
Published: (2024)
Experimental Study of Cold‐formed Steel Frames with Semi‐rigid Floor‐to‐wall Joints
by: Xi Guo, et al.
Published: (2025)
by: Xi Guo, et al.
Published: (2025)
Pre-trained protein language model for codon optimization
by: Pathak, Shashank, et al.
Published: (2024)
by: Pathak, Shashank, et al.
Published: (2024)
Tahakom LLM Guidelines and Recipes: From Pre-training Data to an Arabic LLM
by: AlOtaibi, Areej, et al.
Published: (2025)
by: AlOtaibi, Areej, et al.
Published: (2025)
Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to Signing
by: Jiang, Zifan, et al.
Published: (2025)
by: Jiang, Zifan, et al.
Published: (2025)
Multimodal Autoregressive Pre-training of Large Vision Encoders
by: Fini, Enrico, et al.
Published: (2024)
by: Fini, Enrico, et al.
Published: (2024)
TRecViT: A Recurrent Video Transformer
by: Pătrăucean, Viorica, et al.
Published: (2024)
by: Pătrăucean, Viorica, et al.
Published: (2024)
Explainable Artificial Intelligence Credit Risk Assessment using Machine Learning
by: Shreya, et al.
Published: (2025)
by: Shreya, et al.
Published: (2025)
Beyond the Encoder: Joint Encoder-Decoder Contrastive Pre-Training Improves Dense Prediction
by: Quetin, Sébastien, et al.
Published: (2025)
by: Quetin, Sébastien, et al.
Published: (2025)
Distance-based mutual congestion feature selection with genetic algorithm for high-dimensional medical datasets
by: Nematzadeh, Hossein, et al.
Published: (2024)
by: Nematzadeh, Hossein, et al.
Published: (2024)
Pre-trained Encoder Inference: Revealing Upstream Encoders In Downstream Machine Learning Services
by: Fu, Shaopeng, et al.
Published: (2024)
by: Fu, Shaopeng, et al.
Published: (2024)
Avaliação do pensamento crítico em contexto escolar: uma perspectiva emergente em psicologia
by: Viorica Alich
Published: (2016)
by: Viorica Alich
Published: (2016)
ColBERT-Zero: To Pre-train Or Not To Pre-train ColBERT models
by: Chaffin, Antoine, et al.
Published: (2026)
by: Chaffin, Antoine, et al.
Published: (2026)
Can Distillation Mitigate Backdoor Attacks in Pre-trained Encoders?
by: Han, TIngxu, et al.
Published: (2024)
by: Han, TIngxu, et al.
Published: (2024)
Mutual Information Guided Backdoor Mitigation for Pre-trained Encoders
by: Han, Tingxu, et al.
Published: (2024)
by: Han, Tingxu, et al.
Published: (2024)
Scaling 4D Representations
by: Carreira, João, et al.
Published: (2024)
by: Carreira, João, et al.
Published: (2024)
Rich Insights from Cheap Signals: Efficient Evaluations via Tensor Factorization
by: Polo, Felipe Maia, et al.
Published: (2026)
by: Polo, Felipe Maia, et al.
Published: (2026)
Contrastive Language Video Time Pre-training
by: Liu, Hengyue, et al.
Published: (2024)
by: Liu, Hengyue, et al.
Published: (2024)
Similar Items
-
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
by: Heyward, Joseph, et al.
Published: (2024) -
Memory Consolidation Enables Long-Context Video Understanding
by: Balažević, Ivana, et al.
Published: (2024) -
TAPVid-3D: A Benchmark for Tracking Any Point in 3D
by: Koppula, Skanda, et al.
Published: (2024) -
Perception Test 2025: Challenge Summary and a Unified VQA Extension
by: Heyward, Joseph, et al.
Published: (2026) -
Dynamic Classifier-Free Diffusion Guidance via Online Feedback
by: Papalampidi, Pinelopi, et al.
Published: (2025)