TaleForge: Interactive Multimodal System for Personalized Story Creation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866912453179211776 |
|---|---|
| author | Nguyen, Minh-Loi Le, Quang-Khai Nguyen, Tam V. Tran, Minh-Triet Le, Trung-Nghia |
| author_facet | Nguyen, Minh-Loi Le, Quang-Khai Nguyen, Tam V. Tran, Minh-Triet Le, Trung-Nghia |
| contents | Storytelling is a deeply personal and creative process, yet existing methods often treat users as passive consumers, offering generic plots with limited personalization. This undermines engagement and immersion, especially where individual style or appearance is crucial. We introduce TaleForge, a personalized story-generation system that integrates large language models (LLMs) and text-to-image diffusion to embed users' facial images within both narratives and illustrations. TaleForge features three interconnected modules: Story Generation, where LLMs create narratives and character descriptions from user prompts; Personalized Image Generation, merging users' faces and outfit choices into character illustrations; and Background Generation, creating scene backdrops that incorporate personalized characters. A user study demonstrated heightened engagement and ownership when individuals appeared as protagonists. Participants praised the system's real-time previews and intuitive controls, though they requested finer narrative editing tools. TaleForge advances multimodal storytelling by aligning personalized text and imagery to create immersive, user-centric experiences. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_21832 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | TaleForge: Interactive Multimodal System for Personalized Story Creation Nguyen, Minh-Loi Le, Quang-Khai Nguyen, Tam V. Tran, Minh-Triet Le, Trung-Nghia Computer Vision and Pattern Recognition Storytelling is a deeply personal and creative process, yet existing methods often treat users as passive consumers, offering generic plots with limited personalization. This undermines engagement and immersion, especially where individual style or appearance is crucial. We introduce TaleForge, a personalized story-generation system that integrates large language models (LLMs) and text-to-image diffusion to embed users' facial images within both narratives and illustrations. TaleForge features three interconnected modules: Story Generation, where LLMs create narratives and character descriptions from user prompts; Personalized Image Generation, merging users' faces and outfit choices into character illustrations; and Background Generation, creating scene backdrops that incorporate personalized characters. A user study demonstrated heightened engagement and ownership when individuals appeared as protagonists. Participants praised the system's real-time previews and intuitive controls, though they requested finer narrative editing tools. TaleForge advances multimodal storytelling by aligning personalized text and imagery to create immersive, user-centric experiences. |
| title | TaleForge: Interactive Multimodal System for Personalized Story Creation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2506.21832 |