Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Venkataramanan, Shashanka, Pariza, Valentinos, Salehi, Mohammadreza, Knobel, Lukas, Gidaris, Spyros, Ramzi, Elias, Bursuc, Andrei, Asano, Yuki M. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations
by: Simoncini, Walter, et al.
Published: (2024)
by: Simoncini, Walter, et al.
Published: (2024)
Near, far: Patch-ordering enhances vision foundation models' scene understanding
by: Pariza, Valentinos, et al.
Published: (2024)
by: Pariza, Valentinos, et al.
Published: (2024)
DIP: Unsupervised Dense In-Context Post-training of Visual Representations
by: Sirko-Galouchenko, Sophia, et al.
Published: (2025)
by: Sirko-Galouchenko, Sophia, et al.
Published: (2025)
MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning
by: Salehi, Mohammadreza, et al.
Published: (2025)
by: Salehi, Mohammadreza, et al.
Published: (2025)
Boosting Visual Instruction Tuning with Self-Supervised Guidance
by: Sirko-Galouchenko, Sophia, et al.
Published: (2026)
by: Sirko-Galouchenko, Sophia, et al.
Published: (2026)
MOCA: Self-supervised Representation Learning by Predicting Masked Online Codebook Assignments
by: Gidaris, Spyros, et al.
Published: (2023)
by: Gidaris, Spyros, et al.
Published: (2023)
Learning to Count without Annotations
by: Knobel, Lukas, et al.
Published: (2023)
by: Knobel, Lukas, et al.
Published: (2023)
Is ImageNet worth 1 video? Learning strong image encoders from 1 long unlabelled video
by: Venkataramanan, Shashanka, et al.
Published: (2023)
by: Venkataramanan, Shashanka, et al.
Published: (2023)
VaViM and VaVAM: Autonomous Driving through Video Generative Modeling
by: Bartoccioni, Florent, et al.
Published: (2025)
by: Bartoccioni, Florent, et al.
Published: (2025)
Drive&Segment: Unsupervised Semantic Segmentation of Urban Scenes via Cross-modal Distillation
by: Vobecky, Antonin, et al.
Published: (2022)
by: Vobecky, Antonin, et al.
Published: (2022)
POP-3D: Open-Vocabulary 3D Occupancy Prediction from Images
by: Vobecky, Antonin, et al.
Published: (2024)
by: Vobecky, Antonin, et al.
Published: (2024)
Coevolving Representations in Joint Image-Feature Diffusion
by: Kouzelis, Theodoros, et al.
Published: (2026)
by: Kouzelis, Theodoros, et al.
Published: (2026)
Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction
by: Karypidis, Efstathios, et al.
Published: (2026)
by: Karypidis, Efstathios, et al.
Published: (2026)
OccFeat: Self-supervised Occupancy Feature Prediction for Pretraining BEV Segmentation Networks
by: Sirko-Galouchenko, Sophia, et al.
Published: (2024)
by: Sirko-Galouchenko, Sophia, et al.
Published: (2024)
Self-supervised Learning of Echocardiographic Video Representations via Online Cluster Distillation
by: Mishra, Divyanshu, et al.
Published: (2025)
by: Mishra, Divyanshu, et al.
Published: (2025)
Three Pillars improving Vision Foundation Model Distillation for Lidar
by: Puy, Gilles, et al.
Published: (2023)
by: Puy, Gilles, et al.
Published: (2023)
IPA: An Information-Reconstructive Input Projection Framework for Efficient Foundation Model Adaptation
by: Yin, Yuan, et al.
Published: (2025)
by: Yin, Yuan, et al.
Published: (2025)
3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds
by: Yamada, Ryousuke, et al.
Published: (2025)
by: Yamada, Ryousuke, et al.
Published: (2025)
Advancing Semantic Future Prediction through Multimodal Visual Sequence Transformers
by: Karypidis, Efstathios, et al.
Published: (2025)
by: Karypidis, Efstathios, et al.
Published: (2025)
Matryoshka Representation Learning
by: Kusupati, Aditya, et al.
Published: (2022)
by: Kusupati, Aditya, et al.
Published: (2022)
SPOT: Self-Training with Patch-Order Permutation for Object-Centric Learning with Autoregressive Transformers
by: Kakogeorgiou, Ioannis, et al.
Published: (2023)
by: Kakogeorgiou, Ioannis, et al.
Published: (2023)
Multi-Token Prediction Needs Registers
by: Gerontopoulos, Anastasios, et al.
Published: (2025)
by: Gerontopoulos, Anastasios, et al.
Published: (2025)
DINO-Foresight: Looking into the Future with DINO
by: Karypidis, Efstathios, et al.
Published: (2024)
by: Karypidis, Efstathios, et al.
Published: (2024)
Redefining Normal: A Novel Object-Level Approach for Multi-Object Novelty Detection
by: Salehi, Mohammadreza, et al.
Published: (2024)
by: Salehi, Mohammadreza, et al.
Published: (2024)
Steerable Visual Representations
by: Ruthardt, Jona, et al.
Published: (2026)
by: Ruthardt, Jona, et al.
Published: (2026)
SelEx: Self-Expertise in Fine-Grained Generalized Category Discovery
by: Rastegar, Sarah, et al.
Published: (2024)
by: Rastegar, Sarah, et al.
Published: (2024)
Boosting Generative Image Modeling via Joint Image-Feature Synthesis
by: Kouzelis, Theodoros, et al.
Published: (2025)
by: Kouzelis, Theodoros, et al.
Published: (2025)
Unsupervised Object Localization in the Era of Self-Supervised ViTs: A Survey
by: Siméoni, Oriane, et al.
Published: (2023)
by: Siméoni, Oriane, et al.
Published: (2023)
GeneralAD: Anomaly Detection Across Domains by Attending to Distorted Features
by: Sträter, Luc P. J., et al.
Published: (2024)
by: Sträter, Luc P. J., et al.
Published: (2024)
SIGMA: Sinkhorn-Guided Masked Video Modeling
by: Salehi, Mohammadreza, et al.
Published: (2024)
by: Salehi, Mohammadreza, et al.
Published: (2024)
Valeo4Cast: A Modular Approach to End-to-End Forecasting
by: Xu, Yihong, et al.
Published: (2024)
by: Xu, Yihong, et al.
Published: (2024)
CLIP's Visual Embedding Projector is a Few-shot Cornucopia
by: Fahes, Mohammad, et al.
Published: (2024)
by: Fahes, Mohammad, et al.
Published: (2024)
Vanilla ViT for Automotive Point Cloud Semantic Segmentation
by: Puy, Gilles, et al.
Published: (2026)
by: Puy, Gilles, et al.
Published: (2026)
BIGFix: Bidirectional Image Generation with Token Fixing
by: Besnier, Victor, et al.
Published: (2025)
by: Besnier, Victor, et al.
Published: (2025)
Driving on Registers
by: Kirby, Ellington, et al.
Published: (2026)
by: Kirby, Ellington, et al.
Published: (2026)
Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation
by: Wen, Tiansheng, et al.
Published: (2025)
by: Wen, Tiansheng, et al.
Published: (2025)
Performance monitoring
by: Andrei, Bursuc
Published: (2025)
by: Andrei, Bursuc
Published: (2025)
Frequency-Modulated Visual Restoration for Matryoshka Large Multimodal Models
by: Pan, Qingtao, et al.
Published: (2026)
by: Pan, Qingtao, et al.
Published: (2026)
Matryoshka Gaussian Splatting
by: Guo, Zhilin, et al.
Published: (2026)
by: Guo, Zhilin, et al.
Published: (2026)
Matryoshka Diffusion Models
by: Gu, Jiatao, et al.
Published: (2023)
by: Gu, Jiatao, et al.
Published: (2023)
Similar Items
-
No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations
by: Simoncini, Walter, et al.
Published: (2024) -
Near, far: Patch-ordering enhances vision foundation models' scene understanding
by: Pariza, Valentinos, et al.
Published: (2024) -
DIP: Unsupervised Dense In-Context Post-training of Visual Representations
by: Sirko-Galouchenko, Sophia, et al.
Published: (2025) -
MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning
by: Salehi, Mohammadreza, et al.
Published: (2025) -
Boosting Visual Instruction Tuning with Self-Supervised Guidance
by: Sirko-Galouchenko, Sophia, et al.
Published: (2026)