LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912600578588672 |
|---|---|
| author | Ng, Ho Yin 'Sam' Hsu, Ting-Yao Ramakrishnan, Aashish Anantha Kveton, Branislav Lipka, Nedim Dernoncourt, Franck Lee, Dongwon Yu, Tong Kim, Sungchul Rossi, Ryan A. Huang, Ting-Hao 'Kenneth' |
| author_facet | Ng, Ho Yin 'Sam' Hsu, Ting-Yao Ramakrishnan, Aashish Anantha Kveton, Branislav Lipka, Nedim Dernoncourt, Franck Lee, Dongwon Yu, Tong Kim, Sungchul Rossi, Ryan A. Huang, Ting-Hao 'Kenneth' |
| contents | Figure captions are crucial for helping readers understand and remember a figure's key message. Many models have been developed to generate these captions, helping authors compose better quality captions more easily. Yet, authors almost always need to revise generic AI-generated captions to match their writing style and the domain's style, highlighting the need for personalization. Despite language models' personalization (LaMP) advances, these technologies often focus on text-only settings and rarely address scenarios where both inputs and profiles are multimodal. This paper introduces LaMP-Cap, a dataset for personalized figure caption generation with multimodal figure profiles. For each target figure, LaMP-Cap provides not only the needed inputs, such as figure images, but also up to three other figures from the same document--each with its image, caption, and figure-mentioning paragraphs--as a profile to characterize the context. Experiments with four LLMs show that using profile information consistently helps generate captions closer to the original author-written ones. Ablation studies reveal that images in the profile are more helpful than figure-mentioning paragraphs, highlighting the advantage of using multimodal profiles over text-only ones. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_06561 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles Ng, Ho Yin 'Sam' Hsu, Ting-Yao Ramakrishnan, Aashish Anantha Kveton, Branislav Lipka, Nedim Dernoncourt, Franck Lee, Dongwon Yu, Tong Kim, Sungchul Rossi, Ryan A. Huang, Ting-Hao 'Kenneth' Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition Figure captions are crucial for helping readers understand and remember a figure's key message. Many models have been developed to generate these captions, helping authors compose better quality captions more easily. Yet, authors almost always need to revise generic AI-generated captions to match their writing style and the domain's style, highlighting the need for personalization. Despite language models' personalization (LaMP) advances, these technologies often focus on text-only settings and rarely address scenarios where both inputs and profiles are multimodal. This paper introduces LaMP-Cap, a dataset for personalized figure caption generation with multimodal figure profiles. For each target figure, LaMP-Cap provides not only the needed inputs, such as figure images, but also up to three other figures from the same document--each with its image, caption, and figure-mentioning paragraphs--as a profile to characterize the context. Experiments with four LLMs show that using profile information consistently helps generate captions closer to the original author-written ones. Ablation studies reveal that images in the profile are more helpful than figure-mentioning paragraphs, highlighting the advantage of using multimodal profiles over text-only ones. |
| title | LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles |
| topic | Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2506.06561 |