RealDiffusion: Physics-informed Attention for Multi-character Storybook Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Qi, Chen, Jun, Tsang, Ivor, Dai, Guang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Instructing Text-to-Image Diffusion Models via Classifier-Guided Semantic Optimization
by: Chang, Yuanyuan, et al.
Published: (2025)
by: Chang, Yuanyuan, et al.
Published: (2025)
Self-Assessed Generation: Trustworthy Label Generation for Optical Flow and Stereo Matching in Real-world
by: Ling, Han, et al.
Published: (2024)
by: Ling, Han, et al.
Published: (2024)
StoryState: Agent-Based State Control for Consistent and Editable Storybooks
by: Sarkar, Ayushman, et al.
Published: (2026)
by: Sarkar, Ayushman, et al.
Published: (2026)
Training-Free Dataset Pruning for Instance Segmentation
by: Dai, Yalun, et al.
Published: (2025)
by: Dai, Yalun, et al.
Published: (2025)
Multisize Dataset Condensation
by: He, Yang, et al.
Published: (2024)
by: He, Yang, et al.
Published: (2024)
Boosting Transferability in Vision-Language Attacks via Diversification along the Intersection Region of Adversarial Trajectory
by: Gao, Sensen, et al.
Published: (2024)
by: Gao, Sensen, et al.
Published: (2024)
Input-Aware Sparse Attention for Real-Time Co-Speech Video Generation
by: Lu, Beijia, et al.
Published: (2025)
by: Lu, Beijia, et al.
Published: (2025)
A Lightweight Multi-Scale Attention Framework for Real-Time Spinal Endoscopic Instance Segmentation
by: Lai, Qi, et al.
Published: (2025)
by: Lai, Qi, et al.
Published: (2025)
MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents
by: Xing, Yun, et al.
Published: (2024)
by: Xing, Yun, et al.
Published: (2024)
Multi-Modal Dataset Distillation in the Wild
by: Dang, Zhuohang, et al.
Published: (2025)
by: Dang, Zhuohang, et al.
Published: (2025)
Timestep-Aware Correction for Quantized Diffusion Models
by: Yao, Yuzhe, et al.
Published: (2024)
by: Yao, Yuzhe, et al.
Published: (2024)
Flipped Classroom: Aligning Teacher Attention with Student in Generalized Category Discovery
by: Lin, Haonan, et al.
Published: (2024)
by: Lin, Haonan, et al.
Published: (2024)
Unifying Watermarking via Dimension-Aware Mapping
by: Meng, Jiale, et al.
Published: (2026)
by: Meng, Jiale, et al.
Published: (2026)
IRAD: Implicit Representation-driven Image Resampling against Adversarial Attacks
by: Cao, Yue, et al.
Published: (2023)
by: Cao, Yue, et al.
Published: (2023)
AngleRoCL: Angle-Robust Concept Learning for Physically View-Invariant T2I Adversarial Patches
by: Ji, Wenjun, et al.
Published: (2025)
by: Ji, Wenjun, et al.
Published: (2025)
Tag-Enriched Multi-Attention with Large Language Models for Cross-Domain Sequential Recommendation
by: Wu, Wangyu, et al.
Published: (2025)
by: Wu, Wangyu, et al.
Published: (2025)
Flow-Factory: A Unified Framework for Reinforcement Learning in Flow-Matching Models
by: Ping, Bowen, et al.
Published: (2026)
by: Ping, Bowen, et al.
Published: (2026)
Self-Guidance: Boosting Flow and Diffusion Generation on Their Own
by: Li, Tiancheng, et al.
Published: (2024)
by: Li, Tiancheng, et al.
Published: (2024)
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments
by: Cao, Yue, et al.
Published: (2024)
by: Cao, Yue, et al.
Published: (2024)
Muon-Accelerated Attention Distillation for Real-Time Edge Synthesis via Optimized Latent Diffusion
by: Chen, Weiye, et al.
Published: (2025)
by: Chen, Weiye, et al.
Published: (2025)
LFA-Net: A Lightweight Network with LiteFusion Attention for Retinal Vessel Segmentation
by: Mehmood, Mehwish, et al.
Published: (2025)
by: Mehmood, Mehwish, et al.
Published: (2025)
Generative Diffusion Contrastive Network for Multi-View Clustering
by: Zhu, Jian, et al.
Published: (2025)
by: Zhu, Jian, et al.
Published: (2025)
DiTFastAttn: Attention Compression for Diffusion Transformer Models
by: Yuan, Zhihang, et al.
Published: (2024)
by: Yuan, Zhihang, et al.
Published: (2024)
Multi-task Learning for Real-time Autonomous Driving Leveraging Task-adaptive Attention Generator
by: Choi, Wonhyeok, et al.
Published: (2024)
by: Choi, Wonhyeok, et al.
Published: (2024)
DreamSalon: A Staged Diffusion Framework for Preserving Identity-Context in Editable Face Generation
by: Lin, Haonan, et al.
Published: (2024)
by: Lin, Haonan, et al.
Published: (2024)
Deciphering Emotions in Children Storybooks: A Comparative Analysis of Multimodal LLMs in Educational Applications
by: Asseri, Bushra, et al.
Published: (2025)
by: Asseri, Bushra, et al.
Published: (2025)
LFRA-Net: A Lightweight Focal and Region-Aware Attention Network for Retinal Vessel Segmentatio
by: Mehmood, Mehwish, et al.
Published: (2025)
by: Mehmood, Mehwish, et al.
Published: (2025)
Dataset Color Quantization: A Training-Oriented Framework for Dataset-Level Compression
by: Yu, Chenyue, et al.
Published: (2026)
by: Yu, Chenyue, et al.
Published: (2026)
Personalized Vision via Visual In-Context Learning
by: Jiang, Yuxin, et al.
Published: (2025)
by: Jiang, Yuxin, et al.
Published: (2025)
Blended Latent Diffusion under Attention Control for Real-World Video Editing
by: Liu, Deyin, et al.
Published: (2024)
by: Liu, Deyin, et al.
Published: (2024)
FreqCross: A Multi-Modal Frequency-Spatial Fusion Network for Robust Detection of Stable Diffusion 3.5 Generated Images
by: Yang, Guang
Published: (2025)
by: Yang, Guang
Published: (2025)
Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation
by: Zhao, Min, et al.
Published: (2026)
by: Zhao, Min, et al.
Published: (2026)
Leveraging Bottom-Up and Top-Down Attention for Few-Shot Object Detection
by: Chen, Xianyu, et al.
Published: (2020)
by: Chen, Xianyu, et al.
Published: (2020)
Causal Diffusion Transformers for Generative Modeling
by: Deng, Chaorui, et al.
Published: (2024)
by: Deng, Chaorui, et al.
Published: (2024)
PID: Physics-Informed Diffusion Model for Infrared Image Generation
by: Mao, Fangyuan, et al.
Published: (2024)
by: Mao, Fangyuan, et al.
Published: (2024)
WorldMesh: Generating Navigable Multi-Room 3D Scenes via Mesh-Conditioned Image Diffusion
by: Schneider, Manuel-Andreas, et al.
Published: (2026)
by: Schneider, Manuel-Andreas, et al.
Published: (2026)
Pseudo-D: Informing Multi-View Uncertainty Estimation with Calibrated Neural Training Dynamics
by: Gu, Ang Nan, et al.
Published: (2025)
by: Gu, Ang Nan, et al.
Published: (2025)
Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention
by: Xu, Dejia, et al.
Published: (2024)
by: Xu, Dejia, et al.
Published: (2024)
Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation
by: Wang, Wenjing, et al.
Published: (2023)
by: Wang, Wenjing, et al.
Published: (2023)
Understanding Attention Mechanism in Video Diffusion Models
by: Liu, Bingyan, et al.
Published: (2025)
by: Liu, Bingyan, et al.
Published: (2025)
Similar Items
-
Instructing Text-to-Image Diffusion Models via Classifier-Guided Semantic Optimization
by: Chang, Yuanyuan, et al.
Published: (2025) -
Self-Assessed Generation: Trustworthy Label Generation for Optical Flow and Stereo Matching in Real-world
by: Ling, Han, et al.
Published: (2024) -
StoryState: Agent-Based State Control for Consistent and Editable Storybooks
by: Sarkar, Ayushman, et al.
Published: (2026) -
Training-Free Dataset Pruning for Instance Segmentation
by: Dai, Yalun, et al.
Published: (2025) -
Multisize Dataset Condensation
by: He, Yang, et al.
Published: (2024)