VL-JEPA: Joint Embedding Predictive Architecture for Vision-language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Delong, Shukor, Mustafa, Moutakanni, Theo, Chung, Willy, Yu, Jade, Kasarla, Tejaswi, Bang, Yejin, Bolourchi, Allen, LeCun, Yann, Fung, Pascale
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912868024188928
author Chen, Delong
Shukor, Mustafa
Moutakanni, Theo
Chung, Willy
Yu, Jade
Kasarla, Tejaswi
Bang, Yejin
Bolourchi, Allen
LeCun, Yann
Fung, Pascale
author_facet Chen, Delong
Shukor, Mustafa
Moutakanni, Theo
Chung, Willy
Yu, Jade
Kasarla, Tejaswi
Bang, Yejin
Bolourchi, Allen
LeCun, Yann
Fung, Pascale
contents We introduce VL-JEPA, a vision-language model built on a Joint Embedding Predictive Architecture (JEPA). Instead of autoregressively generating tokens as in classical VLMs, VL-JEPA predicts continuous embeddings of the target texts. By learning in an abstract representation space, the model focuses on task-relevant semantics while abstracting away surface-level linguistic variability. In a strictly controlled comparison against standard token-space VLM training with the same vision encoder and training data, VL-JEPA achieves stronger performance while having 50% fewer trainable parameters. At inference time, a lightweight text decoder is invoked only when needed to translate VL-JEPA predicted embeddings into text. We show that VL-JEPA natively supports selective decoding that reduces the number of decoding operations by 2.85x while maintaining similar performance compared to non-adaptive uniform decoding. Beyond generation, the VL-JEPA's embedding space naturally supports open-vocabulary classification, text-to-video retrieval, and discriminative VQA without any architecture modification. On eight video classification and eight video retrieval datasets, the average performance VL-JEPA surpasses that of CLIP, SigLIP2, and Perception Encoder. At the same time, the model achieves comparable performance as classical VLMs (InstructBLIP, QwenVL) on four VQA datasets: GQA, TallyQA, POPE and POPEv2, despite only having 1.6B parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10942
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
Chen, Delong
Shukor, Mustafa
Moutakanni, Theo
Chung, Willy
Yu, Jade
Kasarla, Tejaswi
Bang, Yejin
Bolourchi, Allen
LeCun, Yann
Fung, Pascale
Computer Vision and Pattern Recognition
We introduce VL-JEPA, a vision-language model built on a Joint Embedding Predictive Architecture (JEPA). Instead of autoregressively generating tokens as in classical VLMs, VL-JEPA predicts continuous embeddings of the target texts. By learning in an abstract representation space, the model focuses on task-relevant semantics while abstracting away surface-level linguistic variability. In a strictly controlled comparison against standard token-space VLM training with the same vision encoder and training data, VL-JEPA achieves stronger performance while having 50% fewer trainable parameters. At inference time, a lightweight text decoder is invoked only when needed to translate VL-JEPA predicted embeddings into text. We show that VL-JEPA natively supports selective decoding that reduces the number of decoding operations by 2.85x while maintaining similar performance compared to non-adaptive uniform decoding. Beyond generation, the VL-JEPA's embedding space naturally supports open-vocabulary classification, text-to-video retrieval, and discriminative VQA without any architecture modification. On eight video classification and eight video retrieval datasets, the average performance VL-JEPA surpasses that of CLIP, SigLIP2, and Perception Encoder. At the same time, the model achieves comparable performance as classical VLMs (InstructBLIP, QwenVL) on four VQA datasets: GQA, TallyQA, POPE and POPEv2, despite only having 1.6B parameters.
title VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.10942