Consistent Story Generation: Unlocking the Potential of Zigzag Sampling
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Mingxiao, Ning, Mang, Moens, Marie-Francine |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models
by: Qu, Tingyu, et al.
Published: (2024)
by: Qu, Tingyu, et al.
Published: (2024)
Towards More Accurate Personalized Image Generation: Addressing Overfitting and Evaluation Bias
by: Li, Mingxiao, et al.
Published: (2025)
by: Li, Mingxiao, et al.
Published: (2025)
Action-based image editing guided by human instructions
by: Trusca, Maria Mihaela, et al.
Published: (2024)
by: Trusca, Maria Mihaela, et al.
Published: (2024)
Alleviating Exposure Bias in Diffusion Models through Sampling with Shifted Time Steps
by: Li, Mingxiao, et al.
Published: (2023)
by: Li, Mingxiao, et al.
Published: (2023)
Animate Your Motion: Turning Still Images into Dynamic Videos
by: Li, Mingxiao, et al.
Published: (2024)
by: Li, Mingxiao, et al.
Published: (2024)
NeuroCine: Decoding Vivid Video Sequences from Human Brain Activties
by: Sun, Jingyuan, et al.
Published: (2024)
by: Sun, Jingyuan, et al.
Published: (2024)
Visually-Aware Context Modeling for News Image Captioning
by: Qu, Tingyu, et al.
Published: (2023)
by: Qu, Tingyu, et al.
Published: (2023)
Introducing Routing Functions to Vision-Language Parameter-Efficient Fine-Tuning with Low-Rank Bottlenecks
by: Qu, Tingyu, et al.
Published: (2024)
by: Qu, Tingyu, et al.
Published: (2024)
DM-Align: Leveraging the Power of Natural Language Instructions to Make Changes to Images
by: Trusca, Maria Mihaela, et al.
Published: (2024)
by: Trusca, Maria Mihaela, et al.
Published: (2024)
Automated Prompt Generation for Creative and Counterfactual Text-to-image Synthesis
by: Jelaca, Aleksa, et al.
Published: (2025)
by: Jelaca, Aleksa, et al.
Published: (2025)
Object-Attribute Binding in Text-to-Image Generation: Evaluation and Control
by: Trusca, Maria Mihaela, et al.
Published: (2024)
by: Trusca, Maria Mihaela, et al.
Published: (2024)
Ctrl-Z Sampling: Scaling Diffusion Sampling with Controlled Random Zigzag Explorations
by: Mao, Shunqi, et al.
Published: (2025)
by: Mao, Shunqi, et al.
Published: (2025)
Spectrum Matching: a Unified Perspective for Superior Diffusability in Latent Diffusion
by: Ning, Mang, et al.
Published: (2026)
by: Ning, Mang, et al.
Published: (2026)
ReDiStory: Region-Disentangled Diffusion for Consistent Visual Story Generation
by: Sarkar, Ayushman, et al.
Published: (2026)
by: Sarkar, Ayushman, et al.
Published: (2026)
StoryMaker: Towards Holistic Consistent Characters in Text-to-image Generation
by: Zhou, Zhengguang, et al.
Published: (2024)
by: Zhou, Zhengguang, et al.
Published: (2024)
Elucidating the Exposure Bias in Diffusion Models
by: Ning, Mang, et al.
Published: (2023)
by: Ning, Mang, et al.
Published: (2023)
$Z^2$-Sampling: Zero-Cost Zigzag Trajectories for Semantic Alignment in Diffusion Models
by: Li, Haosen, et al.
Published: (2026)
by: Li, Haosen, et al.
Published: (2026)
Multimodal Adaptive Inference for Document Image Classification with Anytime Early Exiting
by: Hamed, Omar, et al.
Published: (2024)
by: Hamed, Omar, et al.
Published: (2024)
StoryGPT-V: Large Language Models as Consistent Story Visualizers
by: Shen, Xiaoqian, et al.
Published: (2023)
by: Shen, Xiaoqian, et al.
Published: (2023)
Infinite-Story: A Training-Free Consistent Text-to-Image Generation
by: Park, Jihun, et al.
Published: (2025)
by: Park, Jihun, et al.
Published: (2025)
Rethinking the Zigzag Flattening for Image Reading
by: Zhao, Qingsong, et al.
Published: (2022)
by: Zhao, Qingsong, et al.
Published: (2022)
RiT: Vanilla Diffusion Transformers Suffice in Representation Space
by: Zhang, Le, et al.
Published: (2026)
by: Zhang, Le, et al.
Published: (2026)
Patient4D: Temporally Consistent Patient Body Mesh Recovery from Monocular Operating Room Video
by: Tu, Mingxiao, et al.
Published: (2026)
by: Tu, Mingxiao, et al.
Published: (2026)
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
by: Zhou, Yupeng, et al.
Published: (2024)
by: Zhou, Yupeng, et al.
Published: (2024)
Entropy-Guided k-Guard Sampling for Long-Horizon Autoregressive Video Generation
by: Han, Yizhao, et al.
Published: (2026)
by: Han, Yizhao, et al.
Published: (2026)
SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency
by: Song, Quanjian, et al.
Published: (2025)
by: Song, Quanjian, et al.
Published: (2025)
InfinityStory: Unlimited Video Generation with World Consistency and Character-Aware Shot Transitions
by: Elmoghany, Mohamed, et al.
Published: (2026)
by: Elmoghany, Mohamed, et al.
Published: (2026)
Unseen No More: Unlocking the Potential of CLIP for Generative Zero-shot HOI Detection
by: Guo, Yixin, et al.
Published: (2024)
by: Guo, Yixin, et al.
Published: (2024)
Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration
by: Mo, Sicheng, et al.
Published: (2025)
by: Mo, Sicheng, et al.
Published: (2025)
Optimizing ID Consistency in Multimodal Large Models: Facial Restoration via Alignment, Entanglement, and Disentanglement
by: Dong, Yuran, et al.
Published: (2026)
by: Dong, Yuran, et al.
Published: (2026)
Zigzag Diffusion Sampling: Diffusion Models Can Self-Improve via Self-Reflection
by: Bai, Lichen, et al.
Published: (2024)
by: Bai, Lichen, et al.
Published: (2024)
Scalable and Generalizable Correspondence Pruning via Geometry-Consistent Pre-training
by: Liao, Tangfei, et al.
Published: (2024)
by: Liao, Tangfei, et al.
Published: (2024)
EmoStory: Emotion-Aware Story Generation
by: Yang, Jingyuan, et al.
Published: (2026)
by: Yang, Jingyuan, et al.
Published: (2026)
Representation Learning and Identity Adversarial Training for Facial Behavior Understanding
by: Ning, Mang, et al.
Published: (2024)
by: Ning, Mang, et al.
Published: (2024)
SEED-Story: Multimodal Long Story Generation with Large Language Model
by: Yang, Shuai, et al.
Published: (2024)
by: Yang, Shuai, et al.
Published: (2024)
CLEAR: Unlocking Generative Potential for Degraded Image Understanding in Unified Multimodal Models
by: Hao, Xiangzhao, et al.
Published: (2026)
by: Hao, Xiangzhao, et al.
Published: (2026)
DCTdiff: Intriguing Properties of Image Generative Modeling in the DCT Space
by: Ning, Mang, et al.
Published: (2024)
by: Ning, Mang, et al.
Published: (2024)
Adaptive Visual Conditioning for Semantic Consistency in Diffusion-Based Story Continuation
by: Mousavi, Seyed Mohammad, et al.
Published: (2025)
by: Mousavi, Seyed Mohammad, et al.
Published: (2025)
Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion Models
by: Shen, Fei, et al.
Published: (2024)
by: Shen, Fei, et al.
Published: (2024)
StoryState: Agent-Based State Control for Consistent and Editable Storybooks
by: Sarkar, Ayushman, et al.
Published: (2026)
by: Sarkar, Ayushman, et al.
Published: (2026)
Similar Items
-
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models
by: Qu, Tingyu, et al.
Published: (2024) -
Towards More Accurate Personalized Image Generation: Addressing Overfitting and Evaluation Bias
by: Li, Mingxiao, et al.
Published: (2025) -
Action-based image editing guided by human instructions
by: Trusca, Maria Mihaela, et al.
Published: (2024) -
Alleviating Exposure Bias in Diffusion Models through Sampling with Shifted Time Steps
by: Li, Mingxiao, et al.
Published: (2023) -
Animate Your Motion: Turning Still Images into Dynamic Videos
by: Li, Mingxiao, et al.
Published: (2024)