DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Huiguo, Yang, Huan, Tuo, Zixi, Zhou, Yuan, Wang, Qiuyue, Zhang, Yuhang, Liu, Zeyu, Huang, Wenhao, Chao, Hongyang, Yin, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving Multi-Subject Consistency in Open-Domain Image Generation with Isolation and Reposition Attention
von: He, Huiguo, et al.
Veröffentlicht: (2024)
von: He, Huiguo, et al.
Veröffentlicht: (2024)
A Progressive Evaluation Framework for Multicultural Analysis of Story Visualization
von: Kapuriya, Janak, et al.
Veröffentlicht: (2025)
von: Kapuriya, Janak, et al.
Veröffentlicht: (2025)
ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline Context
von: Zheng, Sixiao, et al.
Veröffentlicht: (2024)
von: Zheng, Sixiao, et al.
Veröffentlicht: (2024)
Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
von: Yan, Xin, et al.
Veröffentlicht: (2024)
von: Yan, Xin, et al.
Veröffentlicht: (2024)
Causal-Story: Local Causal Attention Utilizing Parameter-Efficient Tuning For Visual Story Synthesis
von: Song, Tianyi, et al.
Veröffentlicht: (2023)
von: Song, Tianyi, et al.
Veröffentlicht: (2023)
Hierarchical Knowledge Graphs for Story Understanding in Visual Narratives
von: Chen, Yi-Chun
Veröffentlicht: (2025)
von: Chen, Yi-Chun
Veröffentlicht: (2025)
"You'll Be Alice Adventuring in Wonderland!" Processes, Challenges, and Opportunities of Creating Animated Virtual Reality Stories
von: Yuan, Lin-Ping, et al.
Veröffentlicht: (2025)
von: Yuan, Lin-Ping, et al.
Veröffentlicht: (2025)
AesopAgent: Agent-driven Evolutionary System on Story-to-Video Production
von: Wang, Jiuniu, et al.
Veröffentlicht: (2024)
von: Wang, Jiuniu, et al.
Veröffentlicht: (2024)
Latent Feature-Guided Conditional Diffusion for Generative Image Semantic Communication
von: Chen, Zehao, et al.
Veröffentlicht: (2025)
von: Chen, Zehao, et al.
Veröffentlicht: (2025)
Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs
von: Chen, Jianhao, et al.
Veröffentlicht: (2026)
von: Chen, Jianhao, et al.
Veröffentlicht: (2026)
Enhancing Generalization in Medical Visual Question Answering Tasks via Gradient-Guided Model Perturbation
von: Liu, Gang, et al.
Veröffentlicht: (2024)
von: Liu, Gang, et al.
Veröffentlicht: (2024)
Storybooth: Training-free Multi-Subject Consistency for Improved Visual Storytelling
von: Singh, Jaskirat, et al.
Veröffentlicht: (2025)
von: Singh, Jaskirat, et al.
Veröffentlicht: (2025)
OpenAVS: Training-Free Open-Vocabulary Audio Visual Segmentation with Foundational Models
von: Chen, Shengkai, et al.
Veröffentlicht: (2025)
von: Chen, Shengkai, et al.
Veröffentlicht: (2025)
Vidmento: Creating Video Stories Through Context-Aware Expansion With Generative Video
von: Yeh, Catherine, et al.
Veröffentlicht: (2026)
von: Yeh, Catherine, et al.
Veröffentlicht: (2026)
AudioStory: Generating Long-Form Narrative Audio with Large Language Models
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)
Language-Guided Diffusion Model for Visual Grounding
von: Chen, Sijia, et al.
Veröffentlicht: (2023)
von: Chen, Sijia, et al.
Veröffentlicht: (2023)
Multimodal Semantic Communication for Generative Audio-Driven Video Conferencing
von: Tong, Haonan, et al.
Veröffentlicht: (2024)
von: Tong, Haonan, et al.
Veröffentlicht: (2024)
TSC-PCAC: Voxel Transformer and Sparse Convolution Based Point Cloud Attribute Compression for 3D Broadcasting
von: Guo, Zixi, et al.
Veröffentlicht: (2024)
von: Guo, Zixi, et al.
Veröffentlicht: (2024)
MapStory: Prototyping Editable Map Animations with LLM Agents
von: Gunturu, Aditya, et al.
Veröffentlicht: (2025)
von: Gunturu, Aditya, et al.
Veröffentlicht: (2025)
PixelatedScatter: Arbitrary-level Visual Abstraction for Large-scale Multiclass Scatterplots
von: Guo, Ziheng, et al.
Veröffentlicht: (2025)
von: Guo, Ziheng, et al.
Veröffentlicht: (2025)
FairyGen: Storied Cartoon Video from a Single Child-Drawn Character
von: Zheng, Jiayi, et al.
Veröffentlicht: (2025)
von: Zheng, Jiayi, et al.
Veröffentlicht: (2025)
PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation
von: Wang, Sen, et al.
Veröffentlicht: (2025)
von: Wang, Sen, et al.
Veröffentlicht: (2025)
DreamLLM-3D: Affective Dream Reliving using Large Language Model and 3D Generative AI
von: Liu, Pinyao, et al.
Veröffentlicht: (2025)
von: Liu, Pinyao, et al.
Veröffentlicht: (2025)
Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
von: Wang, Bing, et al.
Veröffentlicht: (2025)
von: Wang, Bing, et al.
Veröffentlicht: (2025)
GiVE: Guiding Visual Encoder to Perceive Overlooked Information
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
DreamFoley: Scalable VLMs for High-Fidelity Video-to-Audio Generation
von: Li, Fu, et al.
Veröffentlicht: (2025)
von: Li, Fu, et al.
Veröffentlicht: (2025)
Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification
von: Qian, Wenhao, et al.
Veröffentlicht: (2025)
von: Qian, Wenhao, et al.
Veröffentlicht: (2025)
CineAGI: Character-Consistent Movie Creation through LLM-Orchestrated Multi-Modal Generation and Cross-Scene Integration
von: Xie, Tianyidan, et al.
Veröffentlicht: (2026)
von: Xie, Tianyidan, et al.
Veröffentlicht: (2026)
Domain-Skewed Federated Learning with Feature Decoupling and Calibration
von: Wang, Huan, et al.
Veröffentlicht: (2026)
von: Wang, Huan, et al.
Veröffentlicht: (2026)
Fast Visual Tracking with Enhanced and Gradient‐Guide Network
von: Dun Cao, et al.
Veröffentlicht: (2024)
von: Dun Cao, et al.
Veröffentlicht: (2024)
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
von: Lv, Zheqi, et al.
Veröffentlicht: (2025)
von: Lv, Zheqi, et al.
Veröffentlicht: (2025)
SyncLipMAE: Contrastive Masked Pretraining for Audio-Visual Talking-Face Representation
von: Ling, Zeyu, et al.
Veröffentlicht: (2025)
von: Ling, Zeyu, et al.
Veröffentlicht: (2025)
Open-Vocabulary Audio-Visual Semantic Segmentation
von: Guo, Ruohao, et al.
Veröffentlicht: (2024)
von: Guo, Ruohao, et al.
Veröffentlicht: (2024)
Language-oriented Semantic Communication for Image Transmission with Fine-Tuned Diffusion Model
von: Wei, Xinfeng, et al.
Veröffentlicht: (2024)
von: Wei, Xinfeng, et al.
Veröffentlicht: (2024)
MindTuner: Cross-Subject Visual Decoding with Visual Fingerprint and Semantic Correction
von: Gong, Zixuan, et al.
Veröffentlicht: (2024)
von: Gong, Zixuan, et al.
Veröffentlicht: (2024)
Towards Automatic Soccer Commentary Generation with Knowledge-Enhanced Visual Reasoning
von: Jin, Zeyu, et al.
Veröffentlicht: (2026)
von: Jin, Zeyu, et al.
Veröffentlicht: (2026)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
von: Su, Fei, et al.
Veröffentlicht: (2026)
von: Su, Fei, et al.
Veröffentlicht: (2026)
Subjective Evaluation of Frame Rate in Bitrate-Constrained Live Streaming
von: He, Jiaqi, et al.
Veröffentlicht: (2026)
von: He, Jiaqi, et al.
Veröffentlicht: (2026)
Uncertainty-driven Sampling for Efficient Pairwise Comparison Subjective Assessment
von: Mohammadi, Shima, et al.
Veröffentlicht: (2024)
von: Mohammadi, Shima, et al.
Veröffentlicht: (2024)
Predictive Sampling for Efficient Pairwise Subjective Image Quality Assessment
von: Mohammadi, Shima, et al.
Veröffentlicht: (2023)
von: Mohammadi, Shima, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Improving Multi-Subject Consistency in Open-Domain Image Generation with Isolation and Reposition Attention
von: He, Huiguo, et al.
Veröffentlicht: (2024) -
A Progressive Evaluation Framework for Multicultural Analysis of Story Visualization
von: Kapuriya, Janak, et al.
Veröffentlicht: (2025) -
ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline Context
von: Zheng, Sixiao, et al.
Veröffentlicht: (2024) -
Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
von: Yan, Xin, et al.
Veröffentlicht: (2024) -
Causal-Story: Local Causal Attention Utilizing Parameter-Efficient Tuning For Visual Story Synthesis
von: Song, Tianyi, et al.
Veröffentlicht: (2023)