Alleviating Distortion in Image Generation via Multi-Resolution Diffusion Models and Time-Dependent Layer Normalization
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Qihao, Zeng, Zhanpeng, He, Ju, Yu, Qihang, Shen, Xiaohui, Chen, Liang-Chieh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FlowTok: Flowing Seamlessly Across Text and Image Tokens
di: He, Ju, et al.
Pubblicazione: (2025)
di: He, Ju, et al.
Pubblicazione: (2025)
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
di: Liu, Qihao, et al.
Pubblicazione: (2025)
di: Liu, Qihao, et al.
Pubblicazione: (2025)
Randomized Autoregressive Visual Generation
di: Yu, Qihang, et al.
Pubblicazione: (2024)
di: Yu, Qihang, et al.
Pubblicazione: (2024)
Frequency-Aware Flow Matching for High-Quality Image Generation
di: Ren, Sucheng, et al.
Pubblicazione: (2026)
di: Ren, Sucheng, et al.
Pubblicazione: (2026)
FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
Autoregressive Image Generation with Masked Bit Modeling
di: Yu, Qihang, et al.
Pubblicazione: (2026)
di: Yu, Qihang, et al.
Pubblicazione: (2026)
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
di: Kim, Dongwon, et al.
Pubblicazione: (2025)
di: Kim, Dongwon, et al.
Pubblicazione: (2025)
Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
di: Ren, Sucheng, et al.
Pubblicazione: (2025)
di: Ren, Sucheng, et al.
Pubblicazione: (2025)
An Image is Worth 32 Tokens for Reconstruction and Generation
di: Yu, Qihang, et al.
Pubblicazione: (2024)
di: Yu, Qihang, et al.
Pubblicazione: (2024)
MaskBit: Embedding-free Image Generation via Bit Tokens
di: Weber, Mark, et al.
Pubblicazione: (2024)
di: Weber, Mark, et al.
Pubblicazione: (2024)
Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers
di: Ren, Sucheng, et al.
Pubblicazione: (2025)
di: Ren, Sucheng, et al.
Pubblicazione: (2025)
A Simple Video Segmenter by Tracking Objects Along Axial Trajectories
di: He, Ju, et al.
Pubblicazione: (2023)
di: He, Ju, et al.
Pubblicazione: (2023)
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
di: Chen, Jieneng, et al.
Pubblicazione: (2024)
di: Chen, Jieneng, et al.
Pubblicazione: (2024)
COCONut: Modernizing COCO Segmentation
di: Deng, Xueqing, et al.
Pubblicazione: (2024)
di: Deng, Xueqing, et al.
Pubblicazione: (2024)
Enhancing Temporal Consistency in Video Editing by Reconstructing Videos with 3D Gaussian Splatting
di: Shin, Inkyu, et al.
Pubblicazione: (2024)
di: Shin, Inkyu, et al.
Pubblicazione: (2024)
COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation
di: Deng, Xueqing, et al.
Pubblicazione: (2025)
di: Deng, Xueqing, et al.
Pubblicazione: (2025)
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
di: Kerssies, Tommie, et al.
Pubblicazione: (2026)
di: Kerssies, Tommie, et al.
Pubblicazione: (2026)
ACCORD: Alleviating Concept Coupling through Dependence Regularization for Text-to-Image Diffusion Personalization
di: Liu, Shizhan, et al.
Pubblicazione: (2025)
di: Liu, Shizhan, et al.
Pubblicazione: (2025)
LayerComposer: Multi-Human Personalized Generation via Layered Canvas
di: Qian, Guocheng Gordon, et al.
Pubblicazione: (2025)
di: Qian, Guocheng Gordon, et al.
Pubblicazione: (2025)
HairDiffusion: Vivid Multi-Colored Hair Editing via Latent Diffusion
di: Zeng, Yu, et al.
Pubblicazione: (2024)
di: Zeng, Yu, et al.
Pubblicazione: (2024)
Dictionary-based Framework for Interpretable and Consistent Object Parsing
di: Zhang, Tiezheng, et al.
Pubblicazione: (2025)
di: Zhang, Tiezheng, et al.
Pubblicazione: (2025)
Breaking Complexity Barriers: High-Resolution Image Restoration with Rank Enhanced Linear Attention
di: Ai, Yuang, et al.
Pubblicazione: (2025)
di: Ai, Yuang, et al.
Pubblicazione: (2025)
DreamLayer: Simultaneous Multi-Layer Generation via Diffusion Mode
di: Huang, Junjia, et al.
Pubblicazione: (2025)
di: Huang, Junjia, et al.
Pubblicazione: (2025)
Towards the Spectral bias Alleviation by Normalizations in Coordinate Networks
di: Cai, Zhicheng, et al.
Pubblicazione: (2024)
di: Cai, Zhicheng, et al.
Pubblicazione: (2024)
Difflare: Removing Image Lens Flare with Latent Diffusion Model
di: Zhou, Tianwen, et al.
Pubblicazione: (2024)
di: Zhou, Tianwen, et al.
Pubblicazione: (2024)
PSDiffusion: Harmonized Multi-Layer Image Generation via Layout and Appearance Alignment
di: Huang, Dingbang, et al.
Pubblicazione: (2025)
di: Huang, Dingbang, et al.
Pubblicazione: (2025)
LayerDiff: Exploring Text-guided Multi-layered Composable Image Synthesis via Layer-Collaborative Diffusion Model
di: Huang, Runhui, et al.
Pubblicazione: (2024)
di: Huang, Runhui, et al.
Pubblicazione: (2024)
RailYolact -- A Yolact Focused on edge for Real-Time Rail Segmentation
di: Qian, Qihao
Pubblicazione: (2024)
di: Qian, Qihao
Pubblicazione: (2024)
Perception-Distortion Balanced Super-Resolution: A Multi-Objective Optimization Perspective
di: Sun, Lingchen, et al.
Pubblicazione: (2023)
di: Sun, Lingchen, et al.
Pubblicazione: (2023)
Multi-Scale Diffusion: Enhancing Spatial Layout in High-Resolution Panoramic Image Generation
di: Zhang, Xiaoyu, et al.
Pubblicazione: (2024)
di: Zhang, Xiaoyu, et al.
Pubblicazione: (2024)
FlashVSR: Towards Real-Time Diffusion-Based Streaming Video Super-Resolution
di: Zhuang, Junhao, et al.
Pubblicazione: (2025)
di: Zhuang, Junhao, et al.
Pubblicazione: (2025)
Explicit Critic Guidance for Aligning Diffusion Models
di: Liang, Zhengyang, et al.
Pubblicazione: (2026)
di: Liang, Zhengyang, et al.
Pubblicazione: (2026)
Perceptual-Distortion Balanced Image Super-Resolution is a Multi-Objective Optimization Problem
di: Zhu, Qiwen, et al.
Pubblicazione: (2024)
di: Zhu, Qiwen, et al.
Pubblicazione: (2024)
Effective Diffusion Transformer Architecture for Image Super-Resolution
di: Cheng, Kun, et al.
Pubblicazione: (2024)
di: Cheng, Kun, et al.
Pubblicazione: (2024)
Evaluating and Predicting Distorted Human Body Parts for Generated Images
di: Ma, Lu, et al.
Pubblicazione: (2025)
di: Ma, Lu, et al.
Pubblicazione: (2025)
Latent Space Super-Resolution for Higher-Resolution Image Generation with Diffusion Models
di: Jeong, Jinho, et al.
Pubblicazione: (2025)
di: Jeong, Jinho, et al.
Pubblicazione: (2025)
Move Anything with Layered Scene Diffusion
di: Ren, Jiawei, et al.
Pubblicazione: (2024)
di: Ren, Jiawei, et al.
Pubblicazione: (2024)
Prior Normality Prompt Transformer for Multi-class Industrial Image Anomaly Detection
di: Yao, Haiming, et al.
Pubblicazione: (2024)
di: Yao, Haiming, et al.
Pubblicazione: (2024)
A Generative Multi-Resolution Pyramid and Normal-Conditioning 3D Cloth Draping
di: Laczkó, Hunor, et al.
Pubblicazione: (2023)
di: Laczkó, Hunor, et al.
Pubblicazione: (2023)
TDEdit: A Unified Diffusion Framework for Text-Drag Guided Image Manipulation
di: Wang, Qihang, et al.
Pubblicazione: (2025)
di: Wang, Qihang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
FlowTok: Flowing Seamlessly Across Text and Image Tokens
di: He, Ju, et al.
Pubblicazione: (2025) -
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
di: Liu, Qihao, et al.
Pubblicazione: (2025) -
Randomized Autoregressive Visual Generation
di: Yu, Qihang, et al.
Pubblicazione: (2024) -
Frequency-Aware Flow Matching for High-Quality Image Generation
di: Ren, Sucheng, et al.
Pubblicazione: (2026) -
FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching
di: Ren, Sucheng, et al.
Pubblicazione: (2024)