Consensus-JEPA: Learned Adaptive Targets from Multi-Mask Predictive Consistency

Fuente: Zenodo
Saved in:
Bibliographic Details
Main Author: Blum, Frederic David
Format: Recurso digital
Published: Zenodo 2026
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866901463391797248
author Blum, Frederic David
author_facet Blum, Frederic David
contents <h2>Title</h2> <p>Consensus-JEPA: Learned Adaptive Targets from Multi-Mask Predictive Consistency</p> <h2>Authors</h2> <p>Frédéric David Blum (freddavidblum@catalystais.com) Catalyst AI — https://catalystais.com/ ORCID: 0009-0009-2487-2974</p> <h2>Upload type</h2> <p>Publication → Preprint</p> <h2>Publication date</h2> <p>2026-03-12</p> <h2>License</h2> <p><strong>All Rights Reserved</strong></p> <p>(This retains full copyright. You can always relax to CC-BY-NC-ND later, but you cannot restrict after publishing under an open license.)</p> <h2>Description (copy-paste this into the Zenodo description field)</h2> <p>This note proposes <strong>Consensus-JEPA</strong>, a minimal extension of the Joint-Embedding Predictive Architecture (I-JEPA) in which the exponential-moving-average (EMA) target encoder is replaced by a <em>learned consensus target module</em> that aggregates predictions across multiple masking configurations.</p> <p>The central question in JEPA design—as of early 2026—is no longer whether to keep the EMA target encoder, but <em>what kind of anchor</em> best replaces it. Existing alternatives include frozen teachers (SALT; Li et al., 2025), pure distributional regularization without any teacher (LeJEPA; Balestriero & LeCun, 2025), and fixed external anchors (GMM-Anchored JEPA, 2026). Consensus-JEPA proposes a different answer: the target should be defined by what remains predictively stable across multiple partial observations of the same input, as judged by an adaptive, learned module.</p> <p><strong>Key contributions:</strong></p> <ol> <li> <p><strong>Leave-one-out consensus target in prediction space.</strong> For each spatial block and each held-out view, the target is computed by aggregating predictions from all remaining masking views. The consensus module operates entirely in prediction space—it never sees raw pixels.</p> </li> <li> <p><strong>Per-block adaptive confidence with pairwise compatibility scoring.</strong> Each view's contribution to the consensus is weighted by (a) a local context-quality score based on compact mask descriptors (visible ratio, spatial distance to context, hole size, connectivity), and (b) a symmetric pairwise compatibility score that verifies mutual consistency between predictions from different views. Symmetry is enforced by construction through absolute-difference and sum inputs. A diversity regularizer with an entropy floor prevents single-view dominance while allowing justified concentration of weight when one view has a genuinely superior context.</p> </li> <li> <p><strong>Consensus resilience diagnostic.</strong> A novel diagnostic (meaningful for K ≥ 4) measures whether the consensus survives partial ablation of contributing views, quantifying how distributed vs. concentrated the consensus is.</p> </li> <li> <p><strong>Integration with SIGReg.</strong> Collapse prevention uses Sketched Isotropic Gaussian Regularization (Balestriero & LeCun, 2025), replacing VICReg-style heuristics with a provably optimal, single-hyperparameter distributional regularizer.</p> </li> <li> <p><strong>Cross-scale extension.</strong> As a secondary contribution, the consensus is extended across spatial scales, requiring target stability under both fine and coarse masking.</p> </li> </ol> <p>The proposal includes a complete specification of modules, exact inputs, training objective with block-aligned cross-view losses, a minimal training algorithm, spectral diagnostics (RankMe, mask-sensitivity index, spectral gap, consensus resilience), and two falsifiable predictions centered on robustness to mask geometry.</p> <p>To the best of our knowledge, among JEPA papers available on arXiv and OpenReview through March 12, 2026, we found no prior method that learns a prediction-space target by leave-one-out aggregation of multiple masked views with per-block confidence weighting and pairwise compatibility scoring.</p> <p><strong>Status:</strong> Architectural hypothesis with falsifiable experimental protocol. Not yet empirically validated.</p> <p><strong>Related work discussed:</strong> I-JEPA (Assran et al., 2023), V-JEPA 2 (Assran et al., 2025), LeJEPA (Balestriero & LeCun, 2025), SALT (Li et al., 2025), VL-JEPA (Chen et al., 2025), LLM-JEPA (Huang et al., 2025), GMM-Anchored JEPA (2026), Causal-JEPA (2026), seq-JEPA (2025), DSeq-JEPA (2025), JEPA as a Neural Tokenizer (2025).</p> <h2>Keywords</h2> <p>JEPA, self-supervised learning, consensus target, multi-view consistency, representation learning, collapse prevention, SIGReg, mask geometry, distribution shift robustness</p> <h2>Related identifiers</h2> <ul> <li>I-JEPA: https://arxiv.org/abs/2301.08243 (isSupplementedBy)</li> <li>LeJEPA: https://arxiv.org/abs/2511.08544 (isSupplementedBy)</li> <li>SALT: https://arxiv.org/abs/2509.24317 (isSupplementedBy)</li> <li>V-JEPA 2: https://arxiv.org/abs/2506.09985 (isSupplementedBy)</li> </ul> <h2>Communities</h2> <p>Machine Learning, Computer Vision, Self-Supervised Learning</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18975567
institution Zenodo
language
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Consensus-JEPA: Learned Adaptive Targets from Multi-Mask Predictive Consistency
Blum, Frederic David
<h2>Title</h2> <p>Consensus-JEPA: Learned Adaptive Targets from Multi-Mask Predictive Consistency</p> <h2>Authors</h2> <p>Frédéric David Blum (freddavidblum@catalystais.com) Catalyst AI — https://catalystais.com/ ORCID: 0009-0009-2487-2974</p> <h2>Upload type</h2> <p>Publication → Preprint</p> <h2>Publication date</h2> <p>2026-03-12</p> <h2>License</h2> <p><strong>All Rights Reserved</strong></p> <p>(This retains full copyright. You can always relax to CC-BY-NC-ND later, but you cannot restrict after publishing under an open license.)</p> <h2>Description (copy-paste this into the Zenodo description field)</h2> <p>This note proposes <strong>Consensus-JEPA</strong>, a minimal extension of the Joint-Embedding Predictive Architecture (I-JEPA) in which the exponential-moving-average (EMA) target encoder is replaced by a <em>learned consensus target module</em> that aggregates predictions across multiple masking configurations.</p> <p>The central question in JEPA design—as of early 2026—is no longer whether to keep the EMA target encoder, but <em>what kind of anchor</em> best replaces it. Existing alternatives include frozen teachers (SALT; Li et al., 2025), pure distributional regularization without any teacher (LeJEPA; Balestriero & LeCun, 2025), and fixed external anchors (GMM-Anchored JEPA, 2026). Consensus-JEPA proposes a different answer: the target should be defined by what remains predictively stable across multiple partial observations of the same input, as judged by an adaptive, learned module.</p> <p><strong>Key contributions:</strong></p> <ol> <li> <p><strong>Leave-one-out consensus target in prediction space.</strong> For each spatial block and each held-out view, the target is computed by aggregating predictions from all remaining masking views. The consensus module operates entirely in prediction space—it never sees raw pixels.</p> </li> <li> <p><strong>Per-block adaptive confidence with pairwise compatibility scoring.</strong> Each view's contribution to the consensus is weighted by (a) a local context-quality score based on compact mask descriptors (visible ratio, spatial distance to context, hole size, connectivity), and (b) a symmetric pairwise compatibility score that verifies mutual consistency between predictions from different views. Symmetry is enforced by construction through absolute-difference and sum inputs. A diversity regularizer with an entropy floor prevents single-view dominance while allowing justified concentration of weight when one view has a genuinely superior context.</p> </li> <li> <p><strong>Consensus resilience diagnostic.</strong> A novel diagnostic (meaningful for K ≥ 4) measures whether the consensus survives partial ablation of contributing views, quantifying how distributed vs. concentrated the consensus is.</p> </li> <li> <p><strong>Integration with SIGReg.</strong> Collapse prevention uses Sketched Isotropic Gaussian Regularization (Balestriero & LeCun, 2025), replacing VICReg-style heuristics with a provably optimal, single-hyperparameter distributional regularizer.</p> </li> <li> <p><strong>Cross-scale extension.</strong> As a secondary contribution, the consensus is extended across spatial scales, requiring target stability under both fine and coarse masking.</p> </li> </ol> <p>The proposal includes a complete specification of modules, exact inputs, training objective with block-aligned cross-view losses, a minimal training algorithm, spectral diagnostics (RankMe, mask-sensitivity index, spectral gap, consensus resilience), and two falsifiable predictions centered on robustness to mask geometry.</p> <p>To the best of our knowledge, among JEPA papers available on arXiv and OpenReview through March 12, 2026, we found no prior method that learns a prediction-space target by leave-one-out aggregation of multiple masked views with per-block confidence weighting and pairwise compatibility scoring.</p> <p><strong>Status:</strong> Architectural hypothesis with falsifiable experimental protocol. Not yet empirically validated.</p> <p><strong>Related work discussed:</strong> I-JEPA (Assran et al., 2023), V-JEPA 2 (Assran et al., 2025), LeJEPA (Balestriero & LeCun, 2025), SALT (Li et al., 2025), VL-JEPA (Chen et al., 2025), LLM-JEPA (Huang et al., 2025), GMM-Anchored JEPA (2026), Causal-JEPA (2026), seq-JEPA (2025), DSeq-JEPA (2025), JEPA as a Neural Tokenizer (2025).</p> <h2>Keywords</h2> <p>JEPA, self-supervised learning, consensus target, multi-view consistency, representation learning, collapse prevention, SIGReg, mask geometry, distribution shift robustness</p> <h2>Related identifiers</h2> <ul> <li>I-JEPA: https://arxiv.org/abs/2301.08243 (isSupplementedBy)</li> <li>LeJEPA: https://arxiv.org/abs/2511.08544 (isSupplementedBy)</li> <li>SALT: https://arxiv.org/abs/2509.24317 (isSupplementedBy)</li> <li>V-JEPA 2: https://arxiv.org/abs/2506.09985 (isSupplementedBy)</li> </ul> <h2>Communities</h2> <p>Machine Learning, Computer Vision, Self-Supervised Learning</p>
title Consensus-JEPA: Learned Adaptive Targets from Multi-Mask Predictive Consistency
url https://doi.org/10.5281/zenodo.18975567