Independent Density Estimation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Jiahao, Cao, Senhao
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917162207150080
author Liu, Jiahao
Cao, Senhao
author_facet Liu, Jiahao
Cao, Senhao
contents Large-scale Vision-Language models have achieved remarkable results in various domains, such as image captioning and conditioned image generation. Nevertheless, these models still encounter difficulties in achieving human-like compositional generalization. In this study, we propose a new method called Independent Density Estimation (IDE) to tackle this challenge. IDE aims to learn the connection between individual words in a sentence and the corresponding features in an image, enabling compositional generalization. We build two models based on the philosophy of IDE. The first one utilizes fully disentangled visual representations as input, and the second leverages a Variational Auto-Encoder to obtain partially disentangled features from raw images. Additionally, we propose an entropy-based compositional inference method to combine predictions of each word in the sentence. Our models exhibit superior generalization to unseen compositions compared to current models when evaluated on various datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10067
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Independent Density Estimation
Liu, Jiahao
Cao, Senhao
Computer Vision and Pattern Recognition
Machine Learning
I.2.6
Large-scale Vision-Language models have achieved remarkable results in various domains, such as image captioning and conditioned image generation. Nevertheless, these models still encounter difficulties in achieving human-like compositional generalization. In this study, we propose a new method called Independent Density Estimation (IDE) to tackle this challenge. IDE aims to learn the connection between individual words in a sentence and the corresponding features in an image, enabling compositional generalization. We build two models based on the philosophy of IDE. The first one utilizes fully disentangled visual representations as input, and the second leverages a Variational Auto-Encoder to obtain partially disentangled features from raw images. Additionally, we propose an entropy-based compositional inference method to combine predictions of each word in the sentence. Our models exhibit superior generalization to unseen compositions compared to current models when evaluated on various datasets.
title Independent Density Estimation
topic Computer Vision and Pattern Recognition
Machine Learning
I.2.6
url https://arxiv.org/abs/2512.10067