Saved in:
| Main Authors: | Chen, Yuzhuo, Ma, Zehua, Wang, Jianhua, Kang, Kai, Yao, Shunyu, Zhang, Weiming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2508.00477 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TAG-WM: Tamper-Aware Generative Image Watermarking via Diffusion Inversion Sensitivity
by: Chen, Yuzhuo, et al.
Published: (2025)
by: Chen, Yuzhuo, et al.
Published: (2025)
Flow of Truth: Proactive Temporal Forensics for Image-to-Video Generation
by: Chen, Yuzhuo, et al.
Published: (2026)
by: Chen, Yuzhuo, et al.
Published: (2026)
CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation
by: Zhang, Hui, et al.
Published: (2024)
by: Zhang, Hui, et al.
Published: (2024)
FreeText: Training-Free Text Rendering in Diffusion Transformers via Attention Localization and Spectral Glyph Injection
by: Zhang, Ruiqiang, et al.
Published: (2026)
by: Zhang, Ruiqiang, et al.
Published: (2026)
I2VWM: Robust Watermarking for Image to Video Generation
by: Wang, Guanjie, et al.
Published: (2025)
by: Wang, Guanjie, et al.
Published: (2025)
Laytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion Transformers
by: Huang, Sida, et al.
Published: (2025)
by: Huang, Sida, et al.
Published: (2025)
Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation
by: Zhang, Haojie, et al.
Published: (2024)
by: Zhang, Haojie, et al.
Published: (2024)
InstanceAssemble: Layout-Aware Image Generation via Instance Assembling Attention
by: Xiang, Qiang, et al.
Published: (2025)
by: Xiang, Qiang, et al.
Published: (2025)
LayoutDETR: Detection Transformer Is a Good Multimodal Layout Designer
by: Yu, Ning, et al.
Published: (2022)
by: Yu, Ning, et al.
Published: (2022)
ConsistCompose: Unified Multimodal Layout Control for Image Composition
by: Shi, Xuanke, et al.
Published: (2025)
by: Shi, Xuanke, et al.
Published: (2025)
Training-Free Layout-to-Image Generation with Marginal Attention Constraints
by: Chen, Huancheng, et al.
Published: (2024)
by: Chen, Huancheng, et al.
Published: (2024)
Manga Generation via Layout-controllable Diffusion
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
Layout2Scene: 3D Semantic Layout Guided Scene Generation via Geometry and Appearance Diffusion Priors
by: Chen, Minglin, et al.
Published: (2025)
by: Chen, Minglin, et al.
Published: (2025)
Retrieval-Augmented Layout Transformer for Content-Aware Layout Generation
by: Horita, Daichi, et al.
Published: (2023)
by: Horita, Daichi, et al.
Published: (2023)
LayoutDiT: Exploring Content-Graphic Balance in Layout Generation with Diffusion Transformer
by: Li, Yu, et al.
Published: (2024)
by: Li, Yu, et al.
Published: (2024)
SKeDA: A Generative Watermarking Framework for Text-to-video Diffusion Models
by: Yang, Yang, et al.
Published: (2026)
by: Yang, Yang, et al.
Published: (2026)
Consistent Image Layout Editing with Diffusion Models
by: Xia, Tao, et al.
Published: (2025)
by: Xia, Tao, et al.
Published: (2025)
Prior Normality Prompt Transformer for Multi-class Industrial Image Anomaly Detection
by: Yao, Haiming, et al.
Published: (2024)
by: Yao, Haiming, et al.
Published: (2024)
Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer
by: Wu, Shuang, et al.
Published: (2024)
by: Wu, Shuang, et al.
Published: (2024)
LayoutAgent: A Vision-Language Agent Guided Compositional Diffusion for Spatial Layout Planning
by: Fan, Zezhong, et al.
Published: (2025)
by: Fan, Zezhong, et al.
Published: (2025)
Layout-Guided Controllable Pathology Image Generation with In-Context Diffusion Transformers
by: Shou, Yuntao, et al.
Published: (2026)
by: Shou, Yuntao, et al.
Published: (2026)
STAY Diffusion: Styled Layout Diffusion Model for Diverse Layout-to-Image Generation
by: Wang, Ruyu, et al.
Published: (2025)
by: Wang, Ruyu, et al.
Published: (2025)
360 Layout Estimation via Orthogonal Planes Disentanglement and Multi-view Geometric Consistency Perception
by: Shen, Zhijie, et al.
Published: (2023)
by: Shen, Zhijie, et al.
Published: (2023)
Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation
by: Ma, Ao, et al.
Published: (2025)
by: Ma, Ao, et al.
Published: (2025)
UniLayDiff: A Unified Diffusion Transformer for Content-Aware Layout Generation
by: Liu, Zeyang, et al.
Published: (2025)
by: Liu, Zeyang, et al.
Published: (2025)
ScalSelect: Scalable Training-Free Multimodal Data Selection for Efficient Visual Instruction Tuning
by: Wu, Changti, et al.
Published: (2026)
by: Wu, Changti, et al.
Published: (2026)
Multitwine: Multi-Object Compositing with Text and Layout Control
by: Tarrés, Gemma Canet, et al.
Published: (2025)
by: Tarrés, Gemma Canet, et al.
Published: (2025)
Multi-Scale Diffusion: Enhancing Spatial Layout in High-Resolution Panoramic Image Generation
by: Zhang, Xiaoyu, et al.
Published: (2024)
by: Zhang, Xiaoyu, et al.
Published: (2024)
M3-AGIQA: Multimodal, Multi-Round, Multi-Aspect AI-Generated Image Quality Assessment
by: Cui, Chuan, et al.
Published: (2025)
by: Cui, Chuan, et al.
Published: (2025)
Relation-Aware Diffusion Model for Controllable Poster Layout Generation
by: Li, Fengheng, et al.
Published: (2023)
by: Li, Fengheng, et al.
Published: (2023)
DLAFormer: An End-to-End Transformer For Document Layout Analysis
by: Wang, Jiawei, et al.
Published: (2024)
by: Wang, Jiawei, et al.
Published: (2024)
MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance
by: Wang, Xierui, et al.
Published: (2024)
by: Wang, Xierui, et al.
Published: (2024)
Griffin: Generative Reference and Layout Guided Image Composition
by: Mikaeili, Aryan, et al.
Published: (2025)
by: Mikaeili, Aryan, et al.
Published: (2025)
Gaussian Shading++: Rethinking the Realistic Deployment Challenge of Performance-Lossless Image Watermark for Diffusion Models
by: Yang, Zijin, et al.
Published: (2025)
by: Yang, Zijin, et al.
Published: (2025)
LayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation
by: Zheng, Guangcong, et al.
Published: (2023)
by: Zheng, Guangcong, et al.
Published: (2023)
FODA-PG for Enhanced Medical Imaging Narrative Generation: Adaptive Differentiation of Normal and Abnormal Attributes
by: Shu, Kai, et al.
Published: (2024)
by: Shu, Kai, et al.
Published: (2024)
Allocentric Perceiver: Disentangling Allocentric Reasoning from Egocentric Visual Priors via Frame Instantiation
by: Wang, Hengyi, et al.
Published: (2026)
by: Wang, Hengyi, et al.
Published: (2026)
VASCAR: Content-Aware Layout Generation via Visual-Aware Self-Correction
by: Zhang, Jiahao, et al.
Published: (2024)
by: Zhang, Jiahao, et al.
Published: (2024)
Gaussian Shading: Provable Performance-Lossless Image Watermarking for Diffusion Models
by: Yang, Zijin, et al.
Published: (2024)
by: Yang, Zijin, et al.
Published: (2024)
Spatial Diffusion for Cell Layout Generation
by: Li, Chen, et al.
Published: (2024)
by: Li, Chen, et al.
Published: (2024)
Similar Items
-
TAG-WM: Tamper-Aware Generative Image Watermarking via Diffusion Inversion Sensitivity
by: Chen, Yuzhuo, et al.
Published: (2025) -
Flow of Truth: Proactive Temporal Forensics for Image-to-Video Generation
by: Chen, Yuzhuo, et al.
Published: (2026) -
CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation
by: Zhang, Hui, et al.
Published: (2024) -
FreeText: Training-Free Text Rendering in Diffusion Transformers via Attention Localization and Spectral Glyph Injection
by: Zhang, Ruiqiang, et al.
Published: (2026) -
I2VWM: Robust Watermarking for Image to Video Generation
by: Wang, Guanjie, et al.
Published: (2025)