Unlocking Compositional Generalization in Continual Few-Shot Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Nguyen-Lam, Phu-Quy, Pham, Phu-Hoa, Minh, Dao Sy Duy, Tran, Chi-Nguyen, Kiet, Huynh Trung, Tran-Thanh, Long
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911693676740608
author Nguyen-Lam, Phu-Quy
Pham, Phu-Hoa
Minh, Dao Sy Duy
Tran, Chi-Nguyen
Kiet, Huynh Trung
Tran-Thanh, Long
author_facet Nguyen-Lam, Phu-Quy
Pham, Phu-Hoa
Minh, Dao Sy Duy
Tran, Chi-Nguyen
Kiet, Huynh Trung
Tran-Thanh, Long
contents Object-centric representations promise a key property for few-shot learning: Rather than treating a scene as a single unit, a model can decompose it into individual object-level parts that can be matched and compared across different concepts. In practice, this potential is rarely realized. Continual learners either collapse scenes into global embeddings, or train with part-level matching objectives that tie representations too closely to seen patterns, leaving them unable to generalize to truly novel concepts. In this paper, we identify this fundamental structural conflict and pioneer a new paradigm that strictly decouples representation learning from compositional inference. Leveraging the inherent patch-level semantic geometry of self-supervised Vision Transformers (ViTs), our framework employs a dual-phase strategy. During training, slot representations are optimized entirely toward holistic class identity, preserving highly generalizable, object-level geometries. At inference, preserved slots are dynamically composed to match novel scenes. We demonstrate that this paradigm offers dual structural benefits: The frozen backbone naturally prevents representation drift, while our lightweight, holistic optimization preserves the features' capacity for novel-concept transfer. Extensive experiments validate this approach, achieving state-of-the-art unseen-concept generalization and minimal forgetting across standard continual learning benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11710
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Unlocking Compositional Generalization in Continual Few-Shot Learning
Nguyen-Lam, Phu-Quy
Pham, Phu-Hoa
Minh, Dao Sy Duy
Tran, Chi-Nguyen
Kiet, Huynh Trung
Tran-Thanh, Long
Machine Learning
Computer Vision and Pattern Recognition
Object-centric representations promise a key property for few-shot learning: Rather than treating a scene as a single unit, a model can decompose it into individual object-level parts that can be matched and compared across different concepts. In practice, this potential is rarely realized. Continual learners either collapse scenes into global embeddings, or train with part-level matching objectives that tie representations too closely to seen patterns, leaving them unable to generalize to truly novel concepts. In this paper, we identify this fundamental structural conflict and pioneer a new paradigm that strictly decouples representation learning from compositional inference. Leveraging the inherent patch-level semantic geometry of self-supervised Vision Transformers (ViTs), our framework employs a dual-phase strategy. During training, slot representations are optimized entirely toward holistic class identity, preserving highly generalizable, object-level geometries. At inference, preserved slots are dynamically composed to match novel scenes. We demonstrate that this paradigm offers dual structural benefits: The frozen backbone naturally prevents representation drift, while our lightweight, holistic optimization preserves the features' capacity for novel-concept transfer. Extensive experiments validate this approach, achieving state-of-the-art unseen-concept generalization and minimal forgetting across standard continual learning benchmarks.
title Unlocking Compositional Generalization in Continual Few-Shot Learning
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.11710