RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909833571074048 |
|---|---|
| author | Oh, Yeongtak Chung, Dohyun Shin, Juhyeon Park, Sangha Barthelemy, Johan Mok, Jisoo Yoon, Sungroh |
| author_facet | Oh, Yeongtak Chung, Dohyun Shin, Juhyeon Park, Sangha Barthelemy, Johan Mok, Jisoo Yoon, Sungroh |
| contents | Recent multi-modal large language models (MLLMs) often struggle to generate personalized image captions, even when trained on high-quality captions. In this work, we observe that such limitations persist in existing post-training-based MLLM personalization methods. Specifically, despite being post-tuned with large-scale caption data through supervised fine-tuning (SFT), these models frequently fail to produce faithful descriptions in real-world scenarios, such as multi-concept image captioning. However, acquiring large-scale, high-quality captions for such complex settings is both costly and difficult. To address the data-centric nature of SFT, we propose a reinforcement learning (RL)-based post-training framework. To the best of our knowledge, this is the first RL-based approach to post-train MLLMs for personalized image captioning. Our method significantly enhances both visual recognition and personalized generation capabilities of MLLMs, and consistently outperforms existing SFT-based baselines, especially in the challenging multi-concept image captioning task. Project page: https://github.com/oyt9306/RePIC |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_18369 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models Oh, Yeongtak Chung, Dohyun Shin, Juhyeon Park, Sangha Barthelemy, Johan Mok, Jisoo Yoon, Sungroh Computer Vision and Pattern Recognition Recent multi-modal large language models (MLLMs) often struggle to generate personalized image captions, even when trained on high-quality captions. In this work, we observe that such limitations persist in existing post-training-based MLLM personalization methods. Specifically, despite being post-tuned with large-scale caption data through supervised fine-tuning (SFT), these models frequently fail to produce faithful descriptions in real-world scenarios, such as multi-concept image captioning. However, acquiring large-scale, high-quality captions for such complex settings is both costly and difficult. To address the data-centric nature of SFT, we propose a reinforcement learning (RL)-based post-training framework. To the best of our knowledge, this is the first RL-based approach to post-train MLLMs for personalized image captioning. Our method significantly enhances both visual recognition and personalized generation capabilities of MLLMs, and consistently outperforms existing SFT-based baselines, especially in the challenging multi-concept image captioning task. Project page: https://github.com/oyt9306/RePIC |
| title | RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2506.18369 |