Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kyaw, Alexander Htet, Sivalingam, Lenin Ravindranath
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909889018724352
author Kyaw, Alexander Htet
Sivalingam, Lenin Ravindranath
author_facet Kyaw, Alexander Htet
Sivalingam, Lenin Ravindranath
contents We present a node-based storytelling system for multimodal content generation. The system represents stories as graphs of nodes that can be expanded, edited, and iteratively refined through direct user edits and natural-language prompts. Each node can integrate text, images, audio, and video, allowing creators to compose multimodal narratives. A task selection agent routes between specialized generative tasks that handle story generation, node structure reasoning, node diagram formatting, and context generation. The interface supports targeted editing of individual nodes, automatic branching for parallel storylines, and node-based iterative refinement. Our results demonstrate that node-based editing supports control over narrative structure and iterative generation of text, images, audio, and video. We report quantitative outcomes on automatic story outline generation and qualitative observations of editing workflows. Finally, we discuss current limitations such as scalability to longer narratives and consistency across multiple nodes, and outline future work toward human-in-the-loop and user-centered creative AI tools.
format Preprint
id arxiv_https___arxiv_org_abs_2511_03227
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video
Kyaw, Alexander Htet
Sivalingam, Lenin Ravindranath
Human-Computer Interaction
Artificial Intelligence
Multimedia
We present a node-based storytelling system for multimodal content generation. The system represents stories as graphs of nodes that can be expanded, edited, and iteratively refined through direct user edits and natural-language prompts. Each node can integrate text, images, audio, and video, allowing creators to compose multimodal narratives. A task selection agent routes between specialized generative tasks that handle story generation, node structure reasoning, node diagram formatting, and context generation. The interface supports targeted editing of individual nodes, automatic branching for parallel storylines, and node-based iterative refinement. Our results demonstrate that node-based editing supports control over narrative structure and iterative generation of text, images, audio, and video. We report quantitative outcomes on automatic story outline generation and qualitative observations of editing workflows. Finally, we discuss current limitations such as scalability to longer narratives and consistency across multiple nodes, and outline future work toward human-in-the-loop and user-centered creative AI tools.
title Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video
topic Human-Computer Interaction
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2511.03227