Saved in:
| Main Authors: | Parthasarathy, Nikhil, Eslami, S. M. Ali, Carreira, João, Hénaff, Olivier J. |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2210.06433 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Layerwise complexity-matched learning yields an improved model of cortical area V2
by: Parthasarathy, Nikhil, et al.
Published: (2023)
by: Parthasarathy, Nikhil, et al.
Published: (2023)
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
by: Garrido, Quentin, et al.
Published: (2025)
by: Garrido, Quentin, et al.
Published: (2025)
Self-supervised vision-langage alignment of deep learning representations for bone X-rays analysis
by: Englebert, Alexandre, et al.
Published: (2024)
by: Englebert, Alexandre, et al.
Published: (2024)
Dilated Convolution with Learnable Spacings makes visual models more aligned with humans: a Grad-CAM study
by: Chamas, Rabih, et al.
Published: (2024)
by: Chamas, Rabih, et al.
Published: (2024)
SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
by: Tschannen, Michael, et al.
Published: (2025)
by: Tschannen, Michael, et al.
Published: (2025)
The role of self-supervised pretraining in differentially private medical image analysis
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
Self-supervised visual learning for analyzing firearms trafficking activities on the Web
by: Konstantakos, Sotirios, et al.
Published: (2023)
by: Konstantakos, Sotirios, et al.
Published: (2023)
Vision Transformer attention alignment with human visual perception in aesthetic object evaluation
by: Carrasco, Miguel, et al.
Published: (2025)
by: Carrasco, Miguel, et al.
Published: (2025)
Advancing human-centric AI for robust X-ray analysis through holistic self-supervised learning
by: Moutakanni, Théo, et al.
Published: (2024)
by: Moutakanni, Théo, et al.
Published: (2024)
The effectiveness of MAE pre-pretraining for billion-scale pretraining
by: Singh, Mannat, et al.
Published: (2023)
by: Singh, Mannat, et al.
Published: (2023)
Dimensions underlying the representational alignment of deep neural networks with humans
by: Mahner, Florian P., et al.
Published: (2024)
by: Mahner, Florian P., et al.
Published: (2024)
Towards aligned body representations in vision models
by: Gizdov, Andrey, et al.
Published: (2025)
by: Gizdov, Andrey, et al.
Published: (2025)
Evaluating alignment between humans and neural network representations in image-based learning tasks
by: Demircan, Can, et al.
Published: (2023)
by: Demircan, Can, et al.
Published: (2023)
Unveiling the Power of Self-supervision for Multi-view Multi-human Association and Tracking
by: Feng, Wei, et al.
Published: (2024)
by: Feng, Wei, et al.
Published: (2024)
Phantom: Subject-consistent video generation via cross-modal alignment
by: Liu, Lijie, et al.
Published: (2025)
by: Liu, Lijie, et al.
Published: (2025)
Predicting upcoming visual features during eye movements yields scene representations aligned with human visual cortex
by: Thorat, Sushrut, et al.
Published: (2025)
by: Thorat, Sushrut, et al.
Published: (2025)
SSTFB: Leveraging self-supervised pretext learning and temporal self-attention with feature branching for real-time video polyp segmentation
by: Xu, Ziang, et al.
Published: (2024)
by: Xu, Ziang, et al.
Published: (2024)
MARS: Paying more attention to visual attributes for text-based person search
by: Ergasti, Alex, et al.
Published: (2024)
by: Ergasti, Alex, et al.
Published: (2024)
Counterfactual contrastive learning: robust representations via causal image synthesis
by: Roschewitz, Melanie, et al.
Published: (2024)
by: Roschewitz, Melanie, et al.
Published: (2024)
A Self-supervised Pressure Map human keypoint Detection Approch: Optimizing Generalization and Computational Efficiency Across Datasets
by: Yu, Chengzhang, et al.
Published: (2024)
by: Yu, Chengzhang, et al.
Published: (2024)
Quantifying the human visual exposome with vision language models
by: Rominger, Christian, et al.
Published: (2026)
by: Rominger, Christian, et al.
Published: (2026)
Breast tumor classification based on self-supervised contrastive learning from ultrasound videos
by: Tang, Yunxin, et al.
Published: (2024)
by: Tang, Yunxin, et al.
Published: (2024)
RO-Bench: Large-scale robustness evaluation of MLLMs with text-driven counterfactual videos
by: Yang, Zixi, et al.
Published: (2025)
by: Yang, Zixi, et al.
Published: (2025)
Improving generalization by mimicking the human visual diet
by: Madan, Spandan, et al.
Published: (2022)
by: Madan, Spandan, et al.
Published: (2022)
Human alignment of neural network representations
by: Muttenthaler, Lukas, et al.
Published: (2022)
by: Muttenthaler, Lukas, et al.
Published: (2022)
Recurrent Video Masked Autoencoders
by: Zoran, Daniel, et al.
Published: (2025)
by: Zoran, Daniel, et al.
Published: (2025)
SuperAnimal pretrained pose estimation models for behavioral analysis
by: Ye, Shaokai, et al.
Published: (2022)
by: Ye, Shaokai, et al.
Published: (2022)
A self-supervised framework for learning whole slide representations
by: Hou, Xinhai, et al.
Published: (2024)
by: Hou, Xinhai, et al.
Published: (2024)
Harnessing small projectors and multiple views for efficient vision pretraining
by: Agrawal, Kumar Krishna, et al.
Published: (2023)
by: Agrawal, Kumar Krishna, et al.
Published: (2023)
Hidden in plain sight: VLMs overlook their visual representations
by: Fu, Stephanie, et al.
Published: (2025)
by: Fu, Stephanie, et al.
Published: (2025)
Masked Modeling for Self-supervised Representation Learning on Vision and Beyond
by: Li, Siyuan, et al.
Published: (2023)
by: Li, Siyuan, et al.
Published: (2023)
Learning from Pattern Completion: Self-supervised Controllable Generation
by: Chen, Zhiqiang, et al.
Published: (2024)
by: Chen, Zhiqiang, et al.
Published: (2024)
Jasmine: Harnessing Diffusion Prior for Self-supervised Depth Estimation
by: Wang, Jiyuan, et al.
Published: (2025)
by: Wang, Jiyuan, et al.
Published: (2025)
Leveraging pretrained RGB denoisers for hyperspectral image restoration
by: Picone, Daniele, et al.
Published: (2026)
by: Picone, Daniele, et al.
Published: (2026)
Point-DAE: Denoising Autoencoders for Self-supervised Point Cloud Learning
by: Zhang, Yabin, et al.
Published: (2022)
by: Zhang, Yabin, et al.
Published: (2022)
AGA: An adaptive group alignment framework for structured medical cross-modal representation learning
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
Car Damage Detection and Patch-to-Patch Self-supervised Image Alignment
by: Chen, Hanxiao
Published: (2024)
by: Chen, Hanxiao
Published: (2024)
A Review on Discriminative Self-supervised Learning Methods in Computer Vision
by: Giakoumoglou, Nikolaos, et al.
Published: (2024)
by: Giakoumoglou, Nikolaos, et al.
Published: (2024)
MAESIL: Masked Autoencoder for Enhanced Self-supervised Medical Image Learning
by: Kim, Kyeonghun, et al.
Published: (2026)
by: Kim, Kyeonghun, et al.
Published: (2026)
LAFS: Landmark-based Facial Self-supervised Learning for Face Recognition
by: Sun, Zhonglin, et al.
Published: (2024)
by: Sun, Zhonglin, et al.
Published: (2024)
Similar Items
-
Layerwise complexity-matched learning yields an improved model of cortical area V2
by: Parthasarathy, Nikhil, et al.
Published: (2023) -
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
by: Garrido, Quentin, et al.
Published: (2025) -
Self-supervised vision-langage alignment of deep learning representations for bone X-rays analysis
by: Englebert, Alexandre, et al.
Published: (2024) -
Dilated Convolution with Learnable Spacings makes visual models more aligned with humans: a Grad-CAM study
by: Chamas, Rabih, et al.
Published: (2024) -
SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
by: Tschannen, Michael, et al.
Published: (2025)