DIP: Unsupervised Dense In-Context Post-training of Visual Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Sirko-Galouchenko, Sophia, Gidaris, Spyros, Vobecky, Antonin, Bursuc, Andrei, Thome, Nicolas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Boosting Visual Instruction Tuning with Self-Supervised Guidance
by: Sirko-Galouchenko, Sophia, et al.
Published: (2026)
by: Sirko-Galouchenko, Sophia, et al.
Published: (2026)
OccFeat: Self-supervised Occupancy Feature Prediction for Pretraining BEV Segmentation Networks
by: Sirko-Galouchenko, Sophia, et al.
Published: (2024)
by: Sirko-Galouchenko, Sophia, et al.
Published: (2024)
Drive&Segment: Unsupervised Semantic Segmentation of Urban Scenes via Cross-modal Distillation
by: Vobecky, Antonin, et al.
Published: (2022)
by: Vobecky, Antonin, et al.
Published: (2022)
POP-3D: Open-Vocabulary 3D Occupancy Prediction from Images
by: Vobecky, Antonin, et al.
Published: (2024)
by: Vobecky, Antonin, et al.
Published: (2024)
MOCA: Self-supervised Representation Learning by Predicting Masked Online Codebook Assignments
by: Gidaris, Spyros, et al.
Published: (2023)
by: Gidaris, Spyros, et al.
Published: (2023)
No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations
by: Simoncini, Walter, et al.
Published: (2024)
by: Simoncini, Walter, et al.
Published: (2024)
Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
by: Venkataramanan, Shashanka, et al.
Published: (2025)
by: Venkataramanan, Shashanka, et al.
Published: (2025)
Test-time Contrastive Concepts for Open-world Semantic Segmentation with Vision-Language Models
by: Wysoczańska, Monika, et al.
Published: (2024)
by: Wysoczańska, Monika, et al.
Published: (2024)
Coevolving Representations in Joint Image-Feature Diffusion
by: Kouzelis, Theodoros, et al.
Published: (2026)
by: Kouzelis, Theodoros, et al.
Published: (2026)
Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction
by: Karypidis, Efstathios, et al.
Published: (2026)
by: Karypidis, Efstathios, et al.
Published: (2026)
Three Pillars improving Vision Foundation Model Distillation for Lidar
by: Puy, Gilles, et al.
Published: (2023)
by: Puy, Gilles, et al.
Published: (2023)
Advancing Semantic Future Prediction through Multimodal Visual Sequence Transformers
by: Karypidis, Efstathios, et al.
Published: (2025)
by: Karypidis, Efstathios, et al.
Published: (2025)
Unsupervised Object Localization in the Era of Self-Supervised ViTs: A Survey
by: Siméoni, Oriane, et al.
Published: (2023)
by: Siméoni, Oriane, et al.
Published: (2023)
SPOT: Self-Training with Patch-Order Permutation for Object-Centric Learning with Autoregressive Transformers
by: Kakogeorgiou, Ioannis, et al.
Published: (2023)
by: Kakogeorgiou, Ioannis, et al.
Published: (2023)
DINO-Foresight: Looking into the Future with DINO
by: Karypidis, Efstathios, et al.
Published: (2024)
by: Karypidis, Efstathios, et al.
Published: (2024)
Multi-Token Prediction Needs Registers
by: Gerontopoulos, Anastasios, et al.
Published: (2025)
by: Gerontopoulos, Anastasios, et al.
Published: (2025)
Boosting Generative Image Modeling via Joint Image-Feature Synthesis
by: Kouzelis, Theodoros, et al.
Published: (2025)
by: Kouzelis, Theodoros, et al.
Published: (2025)
CLIP's Visual Embedding Projector is a Few-shot Cornucopia
by: Fahes, Mohammad, et al.
Published: (2024)
by: Fahes, Mohammad, et al.
Published: (2024)
Valeo4Cast: A Modular Approach to End-to-End Forecasting
by: Xu, Yihong, et al.
Published: (2024)
by: Xu, Yihong, et al.
Published: (2024)
BIGFix: Bidirectional Image Generation with Token Fixing
by: Besnier, Victor, et al.
Published: (2025)
by: Besnier, Victor, et al.
Published: (2025)
Vanilla ViT for Automotive Point Cloud Semantic Segmentation
by: Puy, Gilles, et al.
Published: (2026)
by: Puy, Gilles, et al.
Published: (2026)
3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds
by: Yamada, Ryousuke, et al.
Published: (2025)
by: Yamada, Ryousuke, et al.
Published: (2025)
Driving on Registers
by: Kirby, Ellington, et al.
Published: (2026)
by: Kirby, Ellington, et al.
Published: (2026)
LiDAS: Lighting-driven Dynamic Active Sensing for Nighttime Perception
by: de Moreau, Simon, et al.
Published: (2025)
by: de Moreau, Simon, et al.
Published: (2025)
A Simple Recipe for Language-guided Domain Generalized Segmentation
by: Fahes, Mohammad, et al.
Published: (2023)
by: Fahes, Mohammad, et al.
Published: (2023)
CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation
by: Wysoczańska, Monika, et al.
Published: (2023)
by: Wysoczańska, Monika, et al.
Published: (2023)
Efficient Unsupervised Visual Representation Learning with Explicit Cluster Balancing
by: Metaxas, Ioannis Maniadis, et al.
Published: (2024)
by: Metaxas, Ioannis Maniadis, et al.
Published: (2024)
Efficient Sparse-to-Dense Visual Localization via Compact Gaussian Scene Representation and Accelerated Dense Pose Estimation
by: Li, Zizhuo, et al.
Published: (2026)
by: Li, Zizhuo, et al.
Published: (2026)
DIP-GS: Deep Image Prior For Gaussian Splatting Sparse View Recovery
by: Khatib, Rajaei, et al.
Published: (2025)
by: Khatib, Rajaei, et al.
Published: (2025)
Dense Video Captioning Using Unsupervised Semantic Information
by: Estevam, Valter, et al.
Published: (2021)
by: Estevam, Valter, et al.
Published: (2021)
DenseTRF: Texture-Aware Unsupervised Representation Adaptation for Surgical Scene Dense Prediction
by: Liao, Guiqiu, et al.
Published: (2026)
by: Liao, Guiqiu, et al.
Published: (2026)
Cluster Contrast for Unsupervised Visual Representation Learning
by: Giakoumoglou, Nikolaos, et al.
Published: (2025)
by: Giakoumoglou, Nikolaos, et al.
Published: (2025)
DIP: Diffusion Learning of Inconsistency Pattern for General DeepFake Detection
by: Nie, Fan, et al.
Published: (2024)
by: Nie, Fan, et al.
Published: (2024)
Improving Semantic Uncertainty Quantification in LVLMs with Semantic Gaussian Processes
by: Hoche, Joseph, et al.
Published: (2025)
by: Hoche, Joseph, et al.
Published: (2025)
SuperQuadricOcc: Real-Time Self-Supervised Semantic Occupancy Estimation with Superquadric Volume Rendering
by: Hayes, Seamie, et al.
Published: (2025)
by: Hayes, Seamie, et al.
Published: (2025)
SUGAR: Pre-training 3D Visual Representations for Robotics
by: Chen, Shizhe, et al.
Published: (2024)
by: Chen, Shizhe, et al.
Published: (2024)
LRSLAM: Low-rank Representation of Signed Distance Fields in Dense Visual SLAM System
by: Park, Hongbeen, et al.
Published: (2025)
by: Park, Hongbeen, et al.
Published: (2025)
VaViM and VaVAM: Autonomous Driving through Video Generative Modeling
by: Bartoccioni, Florent, et al.
Published: (2025)
by: Bartoccioni, Florent, et al.
Published: (2025)
Event Camera Data Dense Pre-training
by: Yang, Yan, et al.
Published: (2023)
by: Yang, Yan, et al.
Published: (2023)
BluRef: Unsupervised Image Deblurring with Dense-Matching References
by: Pham, Bang-Dang, et al.
Published: (2026)
by: Pham, Bang-Dang, et al.
Published: (2026)
Similar Items
-
Boosting Visual Instruction Tuning with Self-Supervised Guidance
by: Sirko-Galouchenko, Sophia, et al.
Published: (2026) -
OccFeat: Self-supervised Occupancy Feature Prediction for Pretraining BEV Segmentation Networks
by: Sirko-Galouchenko, Sophia, et al.
Published: (2024) -
Drive&Segment: Unsupervised Semantic Segmentation of Urban Scenes via Cross-modal Distillation
by: Vobecky, Antonin, et al.
Published: (2022) -
POP-3D: Open-Vocabulary 3D Occupancy Prediction from Images
by: Vobecky, Antonin, et al.
Published: (2024) -
MOCA: Self-supervised Representation Learning by Predicting Masked Online Codebook Assignments
by: Gidaris, Spyros, et al.
Published: (2023)