Perception Encoder: The best visual embeddings are not at the output of the network

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bolya, Daniel, Huang, Po-Yao, Sun, Peize, Cho, Jang Hyun, Madotto, Andrea, Wei, Chen, Ma, Tengyu, Zhi, Jiale, Rajasegaran, Jathushan, Rasheed, Hanoona, Wang, Junke, Monteiro, Marco, Xu, Hu, Dong, Shiyu, Ravi, Nikhila, Li, Daniel, Dollár, Piotr, Feichtenhofer, Christoph
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915265460043776
author Bolya, Daniel
Huang, Po-Yao
Sun, Peize
Cho, Jang Hyun
Madotto, Andrea
Wei, Chen
Ma, Tengyu
Zhi, Jiale
Rajasegaran, Jathushan
Rasheed, Hanoona
Wang, Junke
Monteiro, Marco
Xu, Hu
Dong, Shiyu
Ravi, Nikhila
Li, Daniel
Dollár, Piotr
Feichtenhofer, Christoph
author_facet Bolya, Daniel
Huang, Po-Yao
Sun, Peize
Cho, Jang Hyun
Madotto, Andrea
Wei, Chen
Ma, Tengyu
Zhi, Jiale
Rajasegaran, Jathushan
Rasheed, Hanoona
Wang, Junke
Monteiro, Marco
Xu, Hu
Dong, Shiyu
Ravi, Nikhila
Li, Daniel
Dollár, Piotr
Feichtenhofer, Christoph
contents We introduce Perception Encoder (PE), a state-of-the-art vision encoder for image and video understanding trained via simple vision-language learning. Traditionally, vision encoders have relied on a variety of pretraining objectives, each tailored to specific downstream tasks such as classification, captioning, or localization. Surprisingly, after scaling our carefully tuned image pretraining recipe and refining with our robust video data engine, we find that contrastive vision-language training alone can produce strong, general embeddings for all of these downstream tasks. There is only one caveat: these embeddings are hidden within the intermediate layers of the network. To draw them out, we introduce two alignment methods: language alignment for multimodal language modeling, and spatial alignment for dense prediction. Together, our PE family of models achieves best-in-class results on a wide variety of tasks, including (1) zero-shot image and video classification and retrieval, simultaneously obtaining 86.6 average zero-shot ImageNet robustness and 76.9 zero-shot Kinetics-400 video classification; (2) document, image, and video Q&A, enabling 94.6 DocVQA, 80.9 InfographicVQA, and 82.7 PerceptionTest with an 8B LLM; and (3) spatial tasks such as detection, tracking, and depth estimation, setting a new COCO state-of-the-art of 66.0 box mAP. To foster further research, we release our models, code, and novel dataset of synthetically and human-annotated videos: https://github.com/facebookresearch/perception_models
format Preprint
id arxiv_https___arxiv_org_abs_2504_13181
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Perception Encoder: The best visual embeddings are not at the output of the network
Bolya, Daniel
Huang, Po-Yao
Sun, Peize
Cho, Jang Hyun
Madotto, Andrea
Wei, Chen
Ma, Tengyu
Zhi, Jiale
Rajasegaran, Jathushan
Rasheed, Hanoona
Wang, Junke
Monteiro, Marco
Xu, Hu
Dong, Shiyu
Ravi, Nikhila
Li, Daniel
Dollár, Piotr
Feichtenhofer, Christoph
Computer Vision and Pattern Recognition
We introduce Perception Encoder (PE), a state-of-the-art vision encoder for image and video understanding trained via simple vision-language learning. Traditionally, vision encoders have relied on a variety of pretraining objectives, each tailored to specific downstream tasks such as classification, captioning, or localization. Surprisingly, after scaling our carefully tuned image pretraining recipe and refining with our robust video data engine, we find that contrastive vision-language training alone can produce strong, general embeddings for all of these downstream tasks. There is only one caveat: these embeddings are hidden within the intermediate layers of the network. To draw them out, we introduce two alignment methods: language alignment for multimodal language modeling, and spatial alignment for dense prediction. Together, our PE family of models achieves best-in-class results on a wide variety of tasks, including (1) zero-shot image and video classification and retrieval, simultaneously obtaining 86.6 average zero-shot ImageNet robustness and 76.9 zero-shot Kinetics-400 video classification; (2) document, image, and video Q&A, enabling 94.6 DocVQA, 80.9 InfographicVQA, and 82.7 PerceptionTest with an 8B LLM; and (3) spatial tasks such as detection, tracking, and depth estimation, setting a new COCO state-of-the-art of 66.0 box mAP. To foster further research, we release our models, code, and novel dataset of synthetically and human-annotated videos: https://github.com/facebookresearch/perception_models
title Perception Encoder: The best visual embeddings are not at the output of the network
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.13181