DIP: Unsupervised Dense In-Context Post-training of Visual Representations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866908526352269312 |
|---|---|
| author | Sirko-Galouchenko, Sophia Gidaris, Spyros Vobecky, Antonin Bursuc, Andrei Thome, Nicolas |
| author_facet | Sirko-Galouchenko, Sophia Gidaris, Spyros Vobecky, Antonin Bursuc, Andrei Thome, Nicolas |
| contents | We introduce DIP, a novel unsupervised post-training method designed to enhance dense image representations in large-scale pretrained vision encoders for in-context scene understanding. Unlike prior approaches that rely on complex self-distillation architectures, our method trains the vision encoder using pseudo-tasks that explicitly simulate downstream in-context scenarios, inspired by meta-learning principles. To enable post-training on unlabeled data, we propose an automatic mechanism for generating in-context tasks that combines a pretrained diffusion model and the vision encoder itself. DIP is simple, unsupervised, and computationally efficient, requiring less than 9 hours on a single A100 GPU. By learning dense representations through pseudo in-context tasks, it achieves strong performance across a wide variety of downstream real-world in-context scene understanding tasks. It outperforms both the initial vision encoder and prior methods, offering a practical and effective solution for improving dense representations. Code available here: https://github.com/sirkosophia/DIP |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_18463 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | DIP: Unsupervised Dense In-Context Post-training of Visual Representations Sirko-Galouchenko, Sophia Gidaris, Spyros Vobecky, Antonin Bursuc, Andrei Thome, Nicolas Computer Vision and Pattern Recognition We introduce DIP, a novel unsupervised post-training method designed to enhance dense image representations in large-scale pretrained vision encoders for in-context scene understanding. Unlike prior approaches that rely on complex self-distillation architectures, our method trains the vision encoder using pseudo-tasks that explicitly simulate downstream in-context scenarios, inspired by meta-learning principles. To enable post-training on unlabeled data, we propose an automatic mechanism for generating in-context tasks that combines a pretrained diffusion model and the vision encoder itself. DIP is simple, unsupervised, and computationally efficient, requiring less than 9 hours on a single A100 GPU. By learning dense representations through pseudo in-context tasks, it achieves strong performance across a wide variety of downstream real-world in-context scene understanding tasks. It outperforms both the initial vision encoder and prior methods, offering a practical and effective solution for improving dense representations. Code available here: https://github.com/sirkosophia/DIP |
| title | DIP: Unsupervised Dense In-Context Post-training of Visual Representations |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2506.18463 |