DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Xiangteng, Sakai, Shunsuke, Chandhok, Shivam, Beery, Sara, Yuan, Kun, Padoy, Nicolas, Hasegawa, Tatsuhito, Sigal, Leonid
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917350767329280
author He, Xiangteng
Sakai, Shunsuke
Chandhok, Shivam
Beery, Sara
Yuan, Kun
Padoy, Nicolas
Hasegawa, Tatsuhito
Sigal, Leonid
author_facet He, Xiangteng
Sakai, Shunsuke
Chandhok, Shivam
Beery, Sara
Yuan, Kun
Padoy, Nicolas
Hasegawa, Tatsuhito
Sigal, Leonid
contents Recent advances in self-supervised visual representation learning have demonstrated the effectiveness of predictive latent-space objectives for learning transferable features. In particular, Image-based Joint-Embedding Predictive Architecture (I-JEPA) learns representations by predicting latent embeddings of masked target regions from visible context. However, it predicts target regions in parallel and all at once, lacking ability to order predictions meaningfully. Inspired by human visual perception, which attends selectively and progressively from primary to secondary cues, we propose DSeq-JEPA, a Discriminative Sequential Joint-Embedding Predictive Architecture that bridges latent predictive and autoregressive self-supervised learning. Specifically, DSeq-JEPA integrates a discriminatively ordered sequential process with JEPA-style learning objective. This is achieved by (i) identifying primary discriminative regions using an attention-derived saliency map that serves as a proxy for visual importance, and (ii) predicting subsequent regions in discriminative order, inducing a curriculum-like semantic progression from primary to secondary cues in pre-training. Extensive experiments across tasks -- image classification (ImageNet), fine-grained visual categorization (iNaturalist21, CUB, Stanford Cars), detection/segmentation (MS-COCO, ADE20K), and low-level reasoning (CLEVR) -- show that DSeq-JEPA consistently learns more discriminative and generalizable representations compared to I-JEPA variants. Project page: https://github.com/SkyShunsuke/DSeq-JEPA.
format Preprint
id arxiv_https___arxiv_org_abs_2511_17354
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
He, Xiangteng
Sakai, Shunsuke
Chandhok, Shivam
Beery, Sara
Yuan, Kun
Padoy, Nicolas
Hasegawa, Tatsuhito
Sigal, Leonid
Computer Vision and Pattern Recognition
Recent advances in self-supervised visual representation learning have demonstrated the effectiveness of predictive latent-space objectives for learning transferable features. In particular, Image-based Joint-Embedding Predictive Architecture (I-JEPA) learns representations by predicting latent embeddings of masked target regions from visible context. However, it predicts target regions in parallel and all at once, lacking ability to order predictions meaningfully. Inspired by human visual perception, which attends selectively and progressively from primary to secondary cues, we propose DSeq-JEPA, a Discriminative Sequential Joint-Embedding Predictive Architecture that bridges latent predictive and autoregressive self-supervised learning. Specifically, DSeq-JEPA integrates a discriminatively ordered sequential process with JEPA-style learning objective. This is achieved by (i) identifying primary discriminative regions using an attention-derived saliency map that serves as a proxy for visual importance, and (ii) predicting subsequent regions in discriminative order, inducing a curriculum-like semantic progression from primary to secondary cues in pre-training. Extensive experiments across tasks -- image classification (ImageNet), fine-grained visual categorization (iNaturalist21, CUB, Stanford Cars), detection/segmentation (MS-COCO, ADE20K), and low-level reasoning (CLEVR) -- show that DSeq-JEPA consistently learns more discriminative and generalizable representations compared to I-JEPA variants. Project page: https://github.com/SkyShunsuke/DSeq-JEPA.
title DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.17354