Generative Image Layer Decomposition with Visual Effects
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Jinrui, Liu, Qing, Li, Yijun, Kim, Soo Ye, Pakhomov, Daniil, Ren, Mengwei, Zhang, Jianming, Lin, Zhe, Xie, Cihang, Zhou, Yuyin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Unified and Controllable Framework for Layered Image Generation with Visual Effects
von: Yang, Jinrui, et al.
Veröffentlicht: (2026)
von: Yang, Jinrui, et al.
Veröffentlicht: (2026)
LayerFusion: Harmonized Multi-Layer Text-to-Image Generation with Generative Priors
von: Dalva, Yusuf, et al.
Veröffentlicht: (2024)
von: Dalva, Yusuf, et al.
Veröffentlicht: (2024)
Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
von: Zhang, Shilong, et al.
Veröffentlicht: (2025)
von: Zhang, Shilong, et al.
Veröffentlicht: (2025)
EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
von: Ju, Xuan, et al.
Veröffentlicht: (2025)
von: Ju, Xuan, et al.
Veröffentlicht: (2025)
GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
von: Wang, Yuhan, et al.
Veröffentlicht: (2025)
von: Wang, Yuhan, et al.
Veröffentlicht: (2025)
Scaling White-Box Transformers for Vision
von: Yang, Jinrui, et al.
Veröffentlicht: (2024)
von: Yang, Jinrui, et al.
Veröffentlicht: (2024)
$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark
von: Yang, Siwei, et al.
Veröffentlicht: (2025)
von: Yang, Siwei, et al.
Veröffentlicht: (2025)
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
von: Liu, Yanqing, et al.
Veröffentlicht: (2025)
von: Liu, Yanqing, et al.
Veröffentlicht: (2025)
Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training
von: Gao, Yipeng, et al.
Veröffentlicht: (2023)
von: Gao, Yipeng, et al.
Veröffentlicht: (2023)
Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and Reconstruction
von: Cai, Yuanhao, et al.
Veröffentlicht: (2024)
von: Cai, Yuanhao, et al.
Veröffentlicht: (2024)
UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics
von: Chen, Xi, et al.
Veröffentlicht: (2024)
von: Chen, Xi, et al.
Veröffentlicht: (2024)
Object-level Scene Deocclusion
von: Liu, Zhengzhe, et al.
Veröffentlicht: (2024)
von: Liu, Zhengzhe, et al.
Veröffentlicht: (2024)
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
von: Wang, Feng, et al.
Veröffentlicht: (2025)
von: Wang, Feng, et al.
Veröffentlicht: (2025)
How Long Can Unified Multimodal Models Generate Images Reliably? Taming Long-Horizon Interleaved Image Generation via Context Curation
von: Chen, Haoyu, et al.
Veröffentlicht: (2026)
von: Chen, Haoyu, et al.
Veröffentlicht: (2026)
3D-TransUNet for Brain Metastases Segmentation in the BraTS2023 Challenge
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing
von: Hui, Mude, et al.
Veröffentlicht: (2024)
von: Hui, Mude, et al.
Veröffentlicht: (2024)
TurboFill: Adapting Few-step Text-to-image Model for Fast Image Inpainting
von: Xie, Liangbin, et al.
Veröffentlicht: (2025)
von: Xie, Liangbin, et al.
Veröffentlicht: (2025)
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
von: Peng, Cihang, et al.
Veröffentlicht: (2025)
von: Peng, Cihang, et al.
Veröffentlicht: (2025)
Story-Iter: A Training-free Iterative Paradigm for Long Story Visualization
von: Mao, Jiawei, et al.
Veröffentlicht: (2024)
von: Mao, Jiawei, et al.
Veröffentlicht: (2024)
What If We Recaption Billions of Web Images with LLaMA-3?
von: Li, Xianhang, et al.
Veröffentlicht: (2024)
von: Li, Xianhang, et al.
Veröffentlicht: (2024)
Mamba-R: Vision Mamba ALSO Needs Registers
von: Wang, Feng, et al.
Veröffentlicht: (2024)
von: Wang, Feng, et al.
Veröffentlicht: (2024)
M-VAR: Decoupled Scale-wise Autoregressive Modeling for High-Quality Image Generation
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
Adventurer: Optimizing Vision Mamba Architecture Designs for Efficiency
von: Wang, Feng, et al.
Veröffentlicht: (2024)
von: Wang, Feng, et al.
Veröffentlicht: (2024)
CompleteMe: Reference-based Human Image Completion
von: Tsai, Yu-Ju, et al.
Veröffentlicht: (2025)
von: Tsai, Yu-Ju, et al.
Veröffentlicht: (2025)
ARFlow: Autoregressive Flow with Hybrid Linear Attention
von: Hui, Mude, et al.
Veröffentlicht: (2025)
von: Hui, Mude, et al.
Veröffentlicht: (2025)
From Inpainting to Layer Decomposition: Repurposing Generative Inpainting Models for Image Layer Decomposition
von: Chen, Jingxi, et al.
Veröffentlicht: (2025)
von: Chen, Jingxi, et al.
Veröffentlicht: (2025)
Generative Video Propagation
von: Liu, Shaoteng, et al.
Veröffentlicht: (2024)
von: Liu, Shaoteng, et al.
Veröffentlicht: (2024)
Rejuvenating image-GPT as Strong Visual Representation Learners
von: Ren, Sucheng, et al.
Veröffentlicht: (2023)
von: Ren, Sucheng, et al.
Veröffentlicht: (2023)
MedFrameQA: A Multi-Image Medical VQA Benchmark for Clinical Reasoning
von: Yu, Suhao, et al.
Veröffentlicht: (2025)
von: Yu, Suhao, et al.
Veröffentlicht: (2025)
AlbumFill: Album-Guided Reasoning and Retrieval for Personalized Image Completion
von: Tsai, Yu-Ju, et al.
Veröffentlicht: (2026)
von: Tsai, Yu-Ju, et al.
Veröffentlicht: (2026)
Rethinking Global Text Conditioning in Diffusion Transformers
von: Starodubcev, Nikita, et al.
Veröffentlicht: (2026)
von: Starodubcev, Nikita, et al.
Veröffentlicht: (2026)
IMPRINT: Generative Object Compositing by Learning Identity-Preserving Representation
von: Song, Yizhi, et al.
Veröffentlicht: (2024)
von: Song, Yizhi, et al.
Veröffentlicht: (2024)
Thinking Outside the BBox: Unconstrained Generative Object Compositing
von: Tarrés, Gemma Canet, et al.
Veröffentlicht: (2024)
von: Tarrés, Gemma Canet, et al.
Veröffentlicht: (2024)
Omni-MMSI: Toward Identity-attributed Social Interaction Understanding
von: Li, Xinpeng, et al.
Veröffentlicht: (2026)
von: Li, Xinpeng, et al.
Veröffentlicht: (2026)
Qwen-Image-Layered: Towards Inherent Editability via Layer Decomposition
von: Yin, Shengming, et al.
Veröffentlicht: (2025)
von: Yin, Shengming, et al.
Veröffentlicht: (2025)
Filter & Align: Leveraging Human Knowledge to Curate Image-Text Data
von: Zhang, Lei, et al.
Veröffentlicht: (2023)
von: Zhang, Lei, et al.
Veröffentlicht: (2023)
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
A Comprehensive Analysis of Mamba for 3D Volumetric Medical Image Segmentation
von: Wang, Chaohan, et al.
Veröffentlicht: (2025)
von: Wang, Chaohan, et al.
Veröffentlicht: (2025)
MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine
von: Xie, Yunfei, et al.
Veröffentlicht: (2024)
von: Xie, Yunfei, et al.
Veröffentlicht: (2024)
A Semantic Space is Worth 256 Language Descriptions: Make Stronger Segmentation Models with Descriptive Properties
von: Xiao, Junfei, et al.
Veröffentlicht: (2023)
von: Xiao, Junfei, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
A Unified and Controllable Framework for Layered Image Generation with Visual Effects
von: Yang, Jinrui, et al.
Veröffentlicht: (2026) -
LayerFusion: Harmonized Multi-Layer Text-to-Image Generation with Generative Priors
von: Dalva, Yusuf, et al.
Veröffentlicht: (2024) -
Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
von: Zhang, Shilong, et al.
Veröffentlicht: (2025) -
EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
von: Ju, Xuan, et al.
Veröffentlicht: (2025) -
GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
von: Wang, Yuhan, et al.
Veröffentlicht: (2025)