Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Venkataramanan, Shashanka, Pariza, Valentinos, Salehi, Mohammadreza, Knobel, Lukas, Gidaris, Spyros, Ramzi, Elias, Bursuc, Andrei, Asano, Yuki M.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917435214397440
author Venkataramanan, Shashanka
Pariza, Valentinos
Salehi, Mohammadreza
Knobel, Lukas
Gidaris, Spyros
Ramzi, Elias
Bursuc, Andrei
Asano, Yuki M.
author_facet Venkataramanan, Shashanka
Pariza, Valentinos
Salehi, Mohammadreza
Knobel, Lukas
Gidaris, Spyros
Ramzi, Elias
Bursuc, Andrei
Asano, Yuki M.
contents We present Franca (pronounced Fran-ka): free one; the first fully open-source (data, code, weights) vision foundation model that matches and in many cases surpasses the performance of state-of-the-art proprietary models, e.g., DINOv2, CLIP, SigLIPv2, etc. Our approach is grounded in a transparent training pipeline inspired by Web-SSL and uses publicly available data: ImageNet-21K and a subset of ReLAION-2B. Beyond model release, we tackle critical limitations in SSL clustering methods. While modern models rely on assigning image features to large codebooks via clustering algorithms like Sinkhorn-Knopp, they fail to account for the inherent ambiguity in clustering semantics. To address this, we introduce a parameter-efficient, multi-head clustering projector based on nested Matryoshka representations. This design progressively refines features into increasingly fine-grained clusters without increasing the model size, enabling both performance and memory efficiency. Additionally, we propose a novel positional disentanglement strategy that explicitly removes positional biases from dense representations, thereby improving the encoding of semantic content. This leads to consistent gains on several downstream benchmarks, demonstrating the utility of cleaner feature spaces. Our contributions establish a new standard for transparent, high-performance vision models and open a path toward more reproducible and generalizable foundation models for the broader AI community. The code and model checkpoints are available at https://github.com/valeoai/Franca.
format Preprint
id arxiv_https___arxiv_org_abs_2507_14137
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
Venkataramanan, Shashanka
Pariza, Valentinos
Salehi, Mohammadreza
Knobel, Lukas
Gidaris, Spyros
Ramzi, Elias
Bursuc, Andrei
Asano, Yuki M.
Computer Vision and Pattern Recognition
We present Franca (pronounced Fran-ka): free one; the first fully open-source (data, code, weights) vision foundation model that matches and in many cases surpasses the performance of state-of-the-art proprietary models, e.g., DINOv2, CLIP, SigLIPv2, etc. Our approach is grounded in a transparent training pipeline inspired by Web-SSL and uses publicly available data: ImageNet-21K and a subset of ReLAION-2B. Beyond model release, we tackle critical limitations in SSL clustering methods. While modern models rely on assigning image features to large codebooks via clustering algorithms like Sinkhorn-Knopp, they fail to account for the inherent ambiguity in clustering semantics. To address this, we introduce a parameter-efficient, multi-head clustering projector based on nested Matryoshka representations. This design progressively refines features into increasingly fine-grained clusters without increasing the model size, enabling both performance and memory efficiency. Additionally, we propose a novel positional disentanglement strategy that explicitly removes positional biases from dense representations, thereby improving the encoding of semantic content. This leads to consistent gains on several downstream benchmarks, demonstrating the utility of cleaner feature spaces. Our contributions establish a new standard for transparent, high-performance vision models and open a path toward more reproducible and generalizable foundation models for the broader AI community. The code and model checkpoints are available at https://github.com/valeoai/Franca.
title Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.14137