Personalized Image Descriptions from Attention Sequences

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xue, Ruoyu, Le, Hieu, Xu, Jingyi, Mondal, Sounak, Leite, Abe, Zelinsky, Gregory, Hoai, Minh, Samaras, Dimitris
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909947657191424
author Xue, Ruoyu
Le, Hieu
Xu, Jingyi
Mondal, Sounak
Leite, Abe
Zelinsky, Gregory
Hoai, Minh
Samaras, Dimitris
author_facet Xue, Ruoyu
Le, Hieu
Xu, Jingyi
Mondal, Sounak
Leite, Abe
Zelinsky, Gregory
Hoai, Minh
Samaras, Dimitris
contents People can view the same image differently: they focus on different regions, objects, and details in varying orders and describe them in distinct linguistic styles. This leads to substantial variability in image descriptions. However, existing models for personalized image description focus on linguistic style alone, with no prior work leveraging individual viewing patterns. We address this gap by explicitly modeling personalized viewing behavior as a core factor in description generation. Our method, DEPER (DEscription-PERception persona encoder), learns a subject embedding that captures both linguistic style and viewing behavior, guided by an auxiliary attention-prediction task. A lightweight adapter aligns these embeddings with a frozen vision-language model, enabling few-shot personalization without retraining. Across four datasets spanning diverse viewing tasks and both short and detailed descriptions, DEPER achieves a 24% average improvement, showing that modeling personalized attention produces more human-aligned and high-quality descriptions. We posit that understanding how people see helps predict what they say; modeling human diversity in perception can improve both performance and human alignment in multimodal systems.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06662
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Personalized Image Descriptions from Attention Sequences
Xue, Ruoyu
Le, Hieu
Xu, Jingyi
Mondal, Sounak
Leite, Abe
Zelinsky, Gregory
Hoai, Minh
Samaras, Dimitris
Computer Vision and Pattern Recognition
People can view the same image differently: they focus on different regions, objects, and details in varying orders and describe them in distinct linguistic styles. This leads to substantial variability in image descriptions. However, existing models for personalized image description focus on linguistic style alone, with no prior work leveraging individual viewing patterns. We address this gap by explicitly modeling personalized viewing behavior as a core factor in description generation. Our method, DEPER (DEscription-PERception persona encoder), learns a subject embedding that captures both linguistic style and viewing behavior, guided by an auxiliary attention-prediction task. A lightweight adapter aligns these embeddings with a frozen vision-language model, enabling few-shot personalization without retraining. Across four datasets spanning diverse viewing tasks and both short and detailed descriptions, DEPER achieves a 24% average improvement, showing that modeling personalized attention produces more human-aligned and high-quality descriptions. We posit that understanding how people see helps predict what they say; modeling human diversity in perception can improve both performance and human alignment in multimodal systems.
title Personalized Image Descriptions from Attention Sequences
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.06662