Is ImageNet worth 1 video? Learning strong image encoders from 1 long unlabelled video
Fuente:
arXiv
Saved in:
| Main Authors: | Venkataramanan, Shashanka, Rizve, Mamshad Nayeem, Carreira, João, Asano, Yuki M., Avrithis, Yannis |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GAReT: Cross-view Video Geolocalization with Adapters and Auto-Regressive Transformers
by: Pillai, Manu S, et al.
Published: (2024)
by: Pillai, Manu S, et al.
Published: (2024)
FinePseudo: Improving Pseudo-Labelling through Temporal-Alignablity for Semi-Supervised Fine-Grained Action Recognition
by: Dave, Ishan Rajendrakumar, et al.
Published: (2024)
by: Dave, Ishan Rajendrakumar, et al.
Published: (2024)
CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation
by: Zhu, Zixin, et al.
Published: (2025)
by: Zhu, Zixin, et al.
Published: (2025)
Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual Generation
by: Kim, Subin, et al.
Published: (2025)
by: Kim, Subin, et al.
Published: (2025)
MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning
by: Salehi, Mohammadreza, et al.
Published: (2025)
by: Salehi, Mohammadreza, et al.
Published: (2025)
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
by: Swetha, Sirnam, et al.
Published: (2024)
by: Swetha, Sirnam, et al.
Published: (2024)
Open Vocabulary Multi-Label Video Classification
by: Gupta, Rohit, et al.
Published: (2024)
by: Gupta, Rohit, et al.
Published: (2024)
Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
by: Venkataramanan, Shashanka, et al.
Published: (2025)
by: Venkataramanan, Shashanka, et al.
Published: (2025)
Speedrunning ImageNet Diffusion
by: Bhanded, Swayam
Published: (2025)
by: Bhanded, Swayam
Published: (2025)
Unified Alignment Protocol: Making Sense of the Unlabeled Data in New Domains
by: Ahmed, Sabbir, et al.
Published: (2025)
by: Ahmed, Sabbir, et al.
Published: (2025)
Multi-Target Unsupervised Domain Adaptation for Semantic Segmentation without External Data
by: Xu, Yonghao, et al.
Published: (2024)
by: Xu, Yonghao, et al.
Published: (2024)
VidLA: Video-Language Alignment at Scale
by: Rizve, Mamshad Nayeem, et al.
Published: (2024)
by: Rizve, Mamshad Nayeem, et al.
Published: (2024)
CA-Stream: Attention-based pooling for interpretable image recognition
by: Torres, Felipe, et al.
Published: (2024)
by: Torres, Felipe, et al.
Published: (2024)
Toward Errorless Training ImageNet-1k
by: Deng, Bo, et al.
Published: (2025)
by: Deng, Bo, et al.
Published: (2025)
Geometry aware 3D generation from in-the-wild images in ImageNet
by: Shen, Qijia, et al.
Published: (2024)
by: Shen, Qijia, et al.
Published: (2024)
Flaws of ImageNet, Computer Vision's Favourite Dataset
by: Kisel, Nikita, et al.
Published: (2024)
by: Kisel, Nikita, et al.
Published: (2024)
Fine-Grained ImageNet Classification in the Wild
by: Lymperaiou, Maria, et al.
Published: (2023)
by: Lymperaiou, Maria, et al.
Published: (2023)
Accessing Vision Foundation Models via ImageNet-1K
by: Zhang, Yitian, et al.
Published: (2024)
by: Zhang, Yitian, et al.
Published: (2024)
ImageNot: A contrast with ImageNet preserves model rankings
by: Salaudeen, Olawale, et al.
Published: (2024)
by: Salaudeen, Olawale, et al.
Published: (2024)
What Makes ImageNet Look Unlike LAION
by: Shirali, Ali, et al.
Published: (2023)
by: Shirali, Ali, et al.
Published: (2023)
Automated Classification of Model Errors on ImageNet
by: Peychev, Momchil, et al.
Published: (2023)
by: Peychev, Momchil, et al.
Published: (2023)
How far can we go with ImageNet for Text-to-Image generation?
by: Degeorge, L., et al.
Published: (2025)
by: Degeorge, L., et al.
Published: (2025)
ImageNet-OOD: Deciphering Modern Out-of-Distribution Detection Algorithms
by: Yang, William, et al.
Published: (2023)
by: Yang, William, et al.
Published: (2023)
Scaling Up Deep Clustering Methods Beyond ImageNet-1K
by: Adaloglou, Nikolas, et al.
Published: (2024)
by: Adaloglou, Nikolas, et al.
Published: (2024)
An empirical study of the effect of video encoders on Temporal Video Grounding
by: De la Jara, Ignacio M., et al.
Published: (2025)
by: De la Jara, Ignacio M., et al.
Published: (2025)
Can Biases in ImageNet Models Explain Generalization?
by: Gavrikov, Paul, et al.
Published: (2024)
by: Gavrikov, Paul, et al.
Published: (2024)
Infrared Object Detection with Ultra Small ConvNets: Is ImageNet Pretraining Still Useful?
by: Muralidharan, Srikanth, et al.
Published: (2025)
by: Muralidharan, Srikanth, et al.
Published: (2025)
G3DR: Generative 3D Reconstruction in ImageNet
by: Reddy, Pradyumna, et al.
Published: (2024)
by: Reddy, Pradyumna, et al.
Published: (2024)
A Learning Paradigm for Interpretable Gradients
by: Figueroa, Felipe Torres, et al.
Published: (2024)
by: Figueroa, Felipe Torres, et al.
Published: (2024)
Self-supervised video pretraining yields robust and more human-aligned visual representations
by: Parthasarathy, Nikhil, et al.
Published: (2022)
by: Parthasarathy, Nikhil, et al.
Published: (2022)
Composed Image Retrieval for Remote Sensing
by: Psomas, Bill, et al.
Published: (2024)
by: Psomas, Bill, et al.
Published: (2024)
CNN and ViT Efficiency Study on Tiny ImageNet and DermaMNIST Datasets
by: Amangeldi, Aidar, et al.
Published: (2025)
by: Amangeldi, Aidar, et al.
Published: (2025)
Opti-CAM: Optimizing saliency maps for interpretability
by: Zhang, Hanwei, et al.
Published: (2023)
by: Zhang, Hanwei, et al.
Published: (2023)
ConvNet vs Transformer, Supervised vs CLIP: Beyond ImageNet Accuracy
by: Vishniakov, Kirill, et al.
Published: (2023)
by: Vishniakov, Kirill, et al.
Published: (2023)
Composed Image Retrieval for Training-Free Domain Conversion
by: Efthymiadis, Nikos, et al.
Published: (2024)
by: Efthymiadis, Nikos, et al.
Published: (2024)
Transfer Learning from ImageNet for MEG-Based Decoding of Imagined Speech
by: Jhilal, Soufiane, et al.
Published: (2026)
by: Jhilal, Soufiane, et al.
Published: (2026)
Babel-ImageNet: Massively Multilingual Evaluation of Vision-and-Language Representations
by: Geigle, Gregor, et al.
Published: (2023)
by: Geigle, Gregor, et al.
Published: (2023)
Interpreting CLIP: Insights on the Robustness to ImageNet Distribution Shifts
by: Crabbé, Jonathan, et al.
Published: (2023)
by: Crabbé, Jonathan, et al.
Published: (2023)
Online pre-training with long-form videos
by: Kato, Itsuki, et al.
Published: (2024)
by: Kato, Itsuki, et al.
Published: (2024)
Comparative Performance of Finetuned ImageNet Pre-trained Models for Electronic Component Classification
by: Shao, Yidi, et al.
Published: (2025)
by: Shao, Yidi, et al.
Published: (2025)
Similar Items
-
GAReT: Cross-view Video Geolocalization with Adapters and Auto-Regressive Transformers
by: Pillai, Manu S, et al.
Published: (2024) -
FinePseudo: Improving Pseudo-Labelling through Temporal-Alignablity for Semi-Supervised Fine-Grained Action Recognition
by: Dave, Ishan Rajendrakumar, et al.
Published: (2024) -
CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation
by: Zhu, Zixin, et al.
Published: (2025) -
Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual Generation
by: Kim, Subin, et al.
Published: (2025) -
MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning
by: Salehi, Mohammadreza, et al.
Published: (2025)