Text-Conditioned Background Generation for Editable Multi-Layer Documents
Fuente:
arXiv
Saved in:
| Main Authors: | Kang, Taewon, J, Joseph K, Tensmeyer, Chris, Kil, Jihyung, Zhu, Wanrong, Lin, Ming C., Morariu, Vlad I. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AnyDoc: Enhancing Document Generation via Large-Scale HTML/CSS Data Synthesis and Height-Aware Reinforcement Optimization
by: Lin, Jiawei, et al.
Published: (2026)
by: Lin, Jiawei, et al.
Published: (2026)
MiLDEdit: Reasoning-Based Multi-Layer Design Document Editing
by: Lin, Zihao, et al.
Published: (2026)
by: Lin, Zihao, et al.
Published: (2026)
Trajectory-Guided Diffusion for Foreground-Preserving Background Generation in Multi-Layer Documents
by: Kang, Taewon
Published: (2026)
by: Kang, Taewon
Published: (2026)
DocumentCLIP: Linking Figures and Main Body Text in Reflowed Documents
by: Liu, Fuxiao, et al.
Published: (2023)
by: Liu, Fuxiao, et al.
Published: (2023)
NEGATE: Constrained Semantic Guidance for Linguistic Negation in Text-to-Video Diffusion
by: Kang, Taewon, et al.
Published: (2026)
by: Kang, Taewon, et al.
Published: (2026)
II-MMR: Identifying and Improving Multi-modal Multi-hop Reasoning in Visual Question Answering
by: Kil, Jihyung, et al.
Published: (2024)
by: Kil, Jihyung, et al.
Published: (2024)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
by: Fallah, Forouzan, et al.
Published: (2025)
by: Fallah, Forouzan, et al.
Published: (2025)
Character-Centered Dialogue Generation from Scene-Level Prompts
by: Kang, Taewon, et al.
Published: (2025)
by: Kang, Taewon, et al.
Published: (2025)
DocPrune:Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning
by: Choi, Joonmyung, et al.
Published: (2026)
by: Choi, Joonmyung, et al.
Published: (2026)
Scene-Action Prompt Fusion for Coherent Text-to-Video Storytelling
by: Kang, Taewon, et al.
Published: (2025)
by: Kang, Taewon, et al.
Published: (2025)
On Mechanistic Knowledge Localization in Text-to-Image Generative Models
by: Basu, Samyadeep, et al.
Published: (2024)
by: Basu, Samyadeep, et al.
Published: (2024)
VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding
by: Chen, Jian, et al.
Published: (2025)
by: Chen, Jian, et al.
Published: (2025)
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
DCR: Counterfactual Attractor Guidance for Rare Compositional Generation
by: Kang, Taewon, et al.
Published: (2026)
by: Kang, Taewon, et al.
Published: (2026)
COLE: A Hierarchical Generation Framework for Multi-Layered and Editable Graphic Design
by: Jia, Peidong, et al.
Published: (2023)
by: Jia, Peidong, et al.
Published: (2023)
CreatiPoster: Towards Editable and Controllable Multi-Layer Graphic Design Generation
by: Zhang, Zhao, et al.
Published: (2025)
by: Zhang, Zhao, et al.
Published: (2025)
DocSynthv2: A Practical Autoregressive Modeling for Document Generation
by: Biswas, Sanket, et al.
Published: (2024)
by: Biswas, Sanket, et al.
Published: (2024)
Agentic Design Review System
by: Nag, Sayan, et al.
Published: (2025)
by: Nag, Sayan, et al.
Published: (2025)
Qwen-Image-Layered: Towards Inherent Editability via Layer Decomposition
by: Yin, Shengming, et al.
Published: (2025)
by: Yin, Shengming, et al.
Published: (2025)
3D-free meets 3D priors: Novel View Synthesis from a Single Image with Pretrained Diffusion Guidance
by: Kang, Taewon, et al.
Published: (2024)
by: Kang, Taewon, et al.
Published: (2024)
Improving Editability in Image Generation with Layer-wise Memory
by: Kim, Daneul, et al.
Published: (2025)
by: Kim, Daneul, et al.
Published: (2025)
3D Space as a Scratchpad for Editable Text-to-Image Generation
by: Saha, Oindrila, et al.
Published: (2026)
by: Saha, Oindrila, et al.
Published: (2026)
CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers
by: Chen, Weidong, et al.
Published: (2026)
by: Chen, Weidong, et al.
Published: (2026)
Representation Shift: Unifying Token Compression with FlashAttention
by: Choi, Joonmyung, et al.
Published: (2025)
by: Choi, Joonmyung, et al.
Published: (2025)
SVGDreamer++: Advancing Editability and Diversity in Text-Guided SVG Generation
by: Xing, Ximing, et al.
Published: (2024)
by: Xing, Ximing, et al.
Published: (2024)
CompoNeRF: Text-guided Multi-object Compositional NeRF with Editable 3D Scene Layout
by: Bai, Haotian, et al.
Published: (2023)
by: Bai, Haotian, et al.
Published: (2023)
MasterWeaver: Taming Editability and Face Identity for Personalized Text-to-Image Generation
by: Wei, Yuxiang, et al.
Published: (2024)
by: Wei, Yuxiang, et al.
Published: (2024)
EditID: Training-Free Editable ID Customization for Text-to-Image Generation
by: Li, Guandong, et al.
Published: (2025)
by: Li, Guandong, et al.
Published: (2025)
LayerFusion: Harmonized Multi-Layer Text-to-Image Generation with Generative Priors
by: Dalva, Yusuf, et al.
Published: (2024)
by: Dalva, Yusuf, et al.
Published: (2024)
GPT-4V(ision) is a Generalist Web Agent, if Grounded
by: Zheng, Boyuan, et al.
Published: (2024)
by: Zheng, Boyuan, et al.
Published: (2024)
LaDe: Unified Multi-Layered Graphic Media Generation and Decomposition
by: Lungu-Stan, Vlad-Constantin, et al.
Published: (2026)
by: Lungu-Stan, Vlad-Constantin, et al.
Published: (2026)
HART: Human Aligned Reconstruction Transformer
by: Chen, Xiyi, et al.
Published: (2025)
by: Chen, Xiyi, et al.
Published: (2025)
AutoFigure-Edit: Generating Editable Scientific Illustration
by: Lin, Zhen, et al.
Published: (2026)
by: Lin, Zhen, et al.
Published: (2026)
Geometry-Editable and Appearance-Preserving Object Compositon
by: Lin, Jianman, et al.
Published: (2025)
by: Lin, Jianman, et al.
Published: (2025)
Customized Generation Reimagined: Fidelity and Editability Harmonized
by: Jin, Jian, et al.
Published: (2024)
by: Jin, Jian, et al.
Published: (2024)
EditIDv2: Editable ID Customization with Data-Lubricated ID Feature Integration for Text-to-Image Generation
by: Li, Guandong, et al.
Published: (2025)
by: Li, Guandong, et al.
Published: (2025)
FlexDoc: Flexible Document Adaptation through Optimizing both Content and Layout
by: Jiang, Yue, et al.
Published: (2024)
by: Jiang, Yue, et al.
Published: (2024)
Dual-View Visual Contextualization for Web Navigation
by: Kil, Jihyung, et al.
Published: (2024)
by: Kil, Jihyung, et al.
Published: (2024)
EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions
by: Chen, Zhiyuan, et al.
Published: (2024)
by: Chen, Zhiyuan, et al.
Published: (2024)
Are Watermarked Images Editable? SafeMark for Watermark-Preserving Text-Guided Image Editing
by: Wu, Xiaodong, et al.
Published: (2026)
by: Wu, Xiaodong, et al.
Published: (2026)
Similar Items
-
AnyDoc: Enhancing Document Generation via Large-Scale HTML/CSS Data Synthesis and Height-Aware Reinforcement Optimization
by: Lin, Jiawei, et al.
Published: (2026) -
MiLDEdit: Reasoning-Based Multi-Layer Design Document Editing
by: Lin, Zihao, et al.
Published: (2026) -
Trajectory-Guided Diffusion for Foreground-Preserving Background Generation in Multi-Layer Documents
by: Kang, Taewon
Published: (2026) -
DocumentCLIP: Linking Figures and Main Body Text in Reflowed Documents
by: Liu, Fuxiao, et al.
Published: (2023) -
NEGATE: Constrained Semantic Guidance for Linguistic Negation in Text-to-Video Diffusion
by: Kang, Taewon, et al.
Published: (2026)