FORLA: Federated Object-centric Representation Learning with Slot Attention

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liao, Guiqiu, Jogan, Matjaz, Eaton, Eric, Hashimoto, Daniel A.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909868138430464
author Liao, Guiqiu
Jogan, Matjaz
Eaton, Eric
Hashimoto, Daniel A.
author_facet Liao, Guiqiu
Jogan, Matjaz
Eaton, Eric
Hashimoto, Daniel A.
contents Learning efficient visual representations across heterogeneous unlabeled datasets remains a central challenge in federated learning. Effective federated representations require features that are jointly informative across clients while disentangling domain-specific factors without supervision. We introduce FORLA, a novel framework for federated object-centric representation learning and feature adaptation across clients using unsupervised slot attention. At the core of our method is a shared feature adapter, trained collaboratively across clients to adapt features from foundation models, and a shared slot attention module that learns to reconstruct the adapted features. To optimize this adapter, we design a two-branch student-teacher architecture. In each client, a student decoder learns to reconstruct full features from foundation models, while a teacher decoder reconstructs their adapted, low-dimensional counterpart. The shared slot attention module bridges cross-domain learning by aligning object-level representations across clients. Experiments in multiple real-world datasets show that our framework not only outperforms centralized baselines on object discovery but also learns a compact, universal representation that generalizes well across domains. This work highlights federated slot attention as an effective tool for scalable, unsupervised visual representation learning from cross-domain data with distributed concepts.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02964
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FORLA: Federated Object-centric Representation Learning with Slot Attention
Liao, Guiqiu
Jogan, Matjaz
Eaton, Eric
Hashimoto, Daniel A.
Computer Vision and Pattern Recognition
Machine Learning
Learning efficient visual representations across heterogeneous unlabeled datasets remains a central challenge in federated learning. Effective federated representations require features that are jointly informative across clients while disentangling domain-specific factors without supervision. We introduce FORLA, a novel framework for federated object-centric representation learning and feature adaptation across clients using unsupervised slot attention. At the core of our method is a shared feature adapter, trained collaboratively across clients to adapt features from foundation models, and a shared slot attention module that learns to reconstruct the adapted features. To optimize this adapter, we design a two-branch student-teacher architecture. In each client, a student decoder learns to reconstruct full features from foundation models, while a teacher decoder reconstructs their adapted, low-dimensional counterpart. The shared slot attention module bridges cross-domain learning by aligning object-level representations across clients. Experiments in multiple real-world datasets show that our framework not only outperforms centralized baselines on object discovery but also learns a compact, universal representation that generalizes well across domains. This work highlights federated slot attention as an effective tool for scalable, unsupervised visual representation learning from cross-domain data with distributed concepts.
title FORLA: Federated Object-centric Representation Learning with Slot Attention
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2506.02964