PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cho, Jang Hyun, Madotto, Andrea, Mavroudi, Effrosyni, Afouras, Triantafyllos, Nagarajan, Tushar, Maaz, Muhammad, Song, Yale, Ma, Tengyu, Hu, Shuming, Jain, Suyog, Martin, Miguel, Wang, Huiyu, Rasheed, Hanoona, Sun, Peize, Huang, Po-Yao, Bolya, Daniel, Ravi, Nikhila, Jain, Shashank, Stark, Tammy, Moon, Shane, Damavandi, Babak, Lee, Vivian, Westbury, Andrew, Khan, Salman, Krähenbühl, Philipp, Dollár, Piotr, Torresani, Lorenzo, Grauman, Kristen, Feichtenhofer, Christoph
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909702974078976
author Cho, Jang Hyun
Madotto, Andrea
Mavroudi, Effrosyni
Afouras, Triantafyllos
Nagarajan, Tushar
Maaz, Muhammad
Song, Yale
Ma, Tengyu
Hu, Shuming
Jain, Suyog
Martin, Miguel
Wang, Huiyu
Rasheed, Hanoona
Sun, Peize
Huang, Po-Yao
Bolya, Daniel
Ravi, Nikhila
Jain, Shashank
Stark, Tammy
Moon, Shane
Damavandi, Babak
Lee, Vivian
Westbury, Andrew
Khan, Salman
Krähenbühl, Philipp
Dollár, Piotr
Torresani, Lorenzo
Grauman, Kristen
Feichtenhofer, Christoph
author_facet Cho, Jang Hyun
Madotto, Andrea
Mavroudi, Effrosyni
Afouras, Triantafyllos
Nagarajan, Tushar
Maaz, Muhammad
Song, Yale
Ma, Tengyu
Hu, Shuming
Jain, Suyog
Martin, Miguel
Wang, Huiyu
Rasheed, Hanoona
Sun, Peize
Huang, Po-Yao
Bolya, Daniel
Ravi, Nikhila
Jain, Shashank
Stark, Tammy
Moon, Shane
Damavandi, Babak
Lee, Vivian
Westbury, Andrew
Khan, Salman
Krähenbühl, Philipp
Dollár, Piotr
Torresani, Lorenzo
Grauman, Kristen
Feichtenhofer, Christoph
contents Vision-language models are integral to computer vision research, yet many high-performing models remain closed-source, obscuring their data, design and training recipe. The research community has responded by using distillation from black-box models to label training data, achieving strong benchmark results, at the cost of measurable scientific progress. However, without knowing the details of the teacher model and its data sources, scientific progress remains difficult to measure. In this paper, we study building a Perception Language Model (PLM) in a fully open and reproducible framework for transparent research in image and video understanding. We analyze standard training pipelines without distillation from proprietary models and explore large-scale synthetic data to identify critical data gaps, particularly in detailed video understanding. To bridge these gaps, we release 2.8M human-labeled instances of fine-grained video question-answer pairs and spatio-temporally grounded video captions. Additionally, we introduce PLM-VideoBench, a suite for evaluating challenging video understanding tasks focusing on the ability to reason about "what", "where", "when", and "how" of a video. We make our work fully reproducible by providing data, training recipes, code & models. https://github.com/facebookresearch/perception_models
format Preprint
id arxiv_https___arxiv_org_abs_2504_13180
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Cho, Jang Hyun
Madotto, Andrea
Mavroudi, Effrosyni
Afouras, Triantafyllos
Nagarajan, Tushar
Maaz, Muhammad
Song, Yale
Ma, Tengyu
Hu, Shuming
Jain, Suyog
Martin, Miguel
Wang, Huiyu
Rasheed, Hanoona
Sun, Peize
Huang, Po-Yao
Bolya, Daniel
Ravi, Nikhila
Jain, Shashank
Stark, Tammy
Moon, Shane
Damavandi, Babak
Lee, Vivian
Westbury, Andrew
Khan, Salman
Krähenbühl, Philipp
Dollár, Piotr
Torresani, Lorenzo
Grauman, Kristen
Feichtenhofer, Christoph
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Vision-language models are integral to computer vision research, yet many high-performing models remain closed-source, obscuring their data, design and training recipe. The research community has responded by using distillation from black-box models to label training data, achieving strong benchmark results, at the cost of measurable scientific progress. However, without knowing the details of the teacher model and its data sources, scientific progress remains difficult to measure. In this paper, we study building a Perception Language Model (PLM) in a fully open and reproducible framework for transparent research in image and video understanding. We analyze standard training pipelines without distillation from proprietary models and explore large-scale synthetic data to identify critical data gaps, particularly in detailed video understanding. To bridge these gaps, we release 2.8M human-labeled instances of fine-grained video question-answer pairs and spatio-temporally grounded video captions. Additionally, we introduce PLM-VideoBench, a suite for evaluating challenging video understanding tasks focusing on the ability to reason about "what", "where", "when", and "how" of a video. We make our work fully reproducible by providing data, training recipes, code & models. https://github.com/facebookresearch/perception_models
title PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2504.13180