Pixel-Space Post-Training of Latent Diffusion Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Christina, Motwani, Simran, Yu, Matthew, Hou, Ji, Juefei-Xu, Felix, Tsai, Sam, Vajda, Peter, He, Zijian, Wang, Jialiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cache Me if You Can: Accelerating Diffusion Models through Block Caching
von: Wimbauer, Felix, et al.
Veröffentlicht: (2023)
von: Wimbauer, Felix, et al.
Veröffentlicht: (2023)
Llama Learns to Direct: DirectorLLM for Human-Centric Video Generation
von: Song, Kunpeng, et al.
Veröffentlicht: (2024)
von: Song, Kunpeng, et al.
Veröffentlicht: (2024)
Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
von: Ma, Xu, et al.
Veröffentlicht: (2025)
von: Ma, Xu, et al.
Veröffentlicht: (2025)
DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space
von: He, Wenkun, et al.
Veröffentlicht: (2025)
von: He, Wenkun, et al.
Veröffentlicht: (2025)
MoCha: Towards Movie-Grade Talking Character Synthesis
von: Wei, Cong, et al.
Veröffentlicht: (2025)
von: Wei, Cong, et al.
Veröffentlicht: (2025)
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
von: Zhang, Lei, et al.
Veröffentlicht: (2026)
von: Zhang, Lei, et al.
Veröffentlicht: (2026)
PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement
von: Zheng, Haitian, et al.
Veröffentlicht: (2025)
von: Zheng, Haitian, et al.
Veröffentlicht: (2025)
Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
von: Liang, Feng, et al.
Veröffentlicht: (2025)
von: Liang, Feng, et al.
Veröffentlicht: (2025)
LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity
von: Wang, Hongjie, et al.
Veröffentlicht: (2024)
von: Wang, Hongjie, et al.
Veröffentlicht: (2024)
Latent Forcing: Reordering the Diffusion Trajectory for Pixel-Space Image Generation
von: Baade, Alan, et al.
Veröffentlicht: (2026)
von: Baade, Alan, et al.
Veröffentlicht: (2026)
Beyond Pixels: Semantic-aware Typographic Attack for Geo-Privacy Protection
von: Zhu, Jiayi, et al.
Veröffentlicht: (2025)
von: Zhu, Jiayi, et al.
Veröffentlicht: (2025)
Exploring MLLM-Diffusion Information Transfer with MetaCanvas
von: Lin, Han, et al.
Veröffentlicht: (2025)
von: Lin, Han, et al.
Veröffentlicht: (2025)
Transforming Weather Data from Pixel to Latent Space
von: Zhao, Sijie, et al.
Veröffentlicht: (2025)
von: Zhao, Sijie, et al.
Veröffentlicht: (2025)
FBNetV5: Neural Architecture Search for Multiple Tasks in One Run
von: Wu, Bichen, et al.
Veröffentlicht: (2021)
von: Wu, Bichen, et al.
Veröffentlicht: (2021)
DiP: Taming Diffusion Models in Pixel Space
von: Chen, Zhennan, et al.
Veröffentlicht: (2025)
von: Chen, Zhennan, et al.
Veröffentlicht: (2025)
HyperDiT: Hyper-Connected Transformers for High-Fidelity Pixel-Space Diffusion
von: He, Yu, et al.
Veröffentlicht: (2026)
von: He, Yu, et al.
Veröffentlicht: (2026)
Registers Matter for Pixel-Space Diffusion Transformers
von: Starodubcev, Nikita, et al.
Veröffentlicht: (2026)
von: Starodubcev, Nikita, et al.
Veröffentlicht: (2026)
Probability Density Geodesics in Image Diffusion Latent Space
von: Yu, Qingtao, et al.
Veröffentlicht: (2025)
von: Yu, Qingtao, et al.
Veröffentlicht: (2025)
Lesion-Aware Post-Training of Latent Diffusion Models for Synthesizing Diffusion MRI from CT Perfusion
von: Lee, Junhyeok, et al.
Veröffentlicht: (2025)
von: Lee, Junhyeok, et al.
Veröffentlicht: (2025)
PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
von: Lu, Yifan, et al.
Veröffentlicht: (2026)
von: Lu, Yifan, et al.
Veröffentlicht: (2026)
Light the Night: A Multi-Condition Diffusion Framework for Unpaired Low-Light Enhancement in Autonomous Driving
von: Li, Jinlong, et al.
Veröffentlicht: (2024)
von: Li, Jinlong, et al.
Veröffentlicht: (2024)
Novel View Synthesis with Pixel-Space Diffusion Models
von: Elata, Noam, et al.
Veröffentlicht: (2024)
von: Elata, Noam, et al.
Veröffentlicht: (2024)
Transfer between Modalities with MetaQueries
von: Pan, Xichen, et al.
Veröffentlicht: (2025)
von: Pan, Xichen, et al.
Veröffentlicht: (2025)
Populate-A-Scene: Affordance-Aware Human Video Generation
von: Shan, Mengyi, et al.
Veröffentlicht: (2025)
von: Shan, Mengyi, et al.
Veröffentlicht: (2025)
PixelDiT: Pixel Diffusion Transformers for Image Generation
von: Yu, Yongsheng, et al.
Veröffentlicht: (2025)
von: Yu, Yongsheng, et al.
Veröffentlicht: (2025)
The Thinking Pixel: Recursive Sparse Reasoning in Multimodal Diffusion Latents
von: Sun, Yuwei, et al.
Veröffentlicht: (2026)
von: Sun, Yuwei, et al.
Veröffentlicht: (2026)
Your Latent Mask is Wrong: Pixel-Equivalent Latent Compositing for Diffusion Models
von: Bradbury, Rowan, et al.
Veröffentlicht: (2025)
von: Bradbury, Rowan, et al.
Veröffentlicht: (2025)
Joint Geometry-Appearance Human Reconstruction in a Unified Latent Space via Bridge Diffusion
von: Tang, Yingzhi, et al.
Veröffentlicht: (2026)
von: Tang, Yingzhi, et al.
Veröffentlicht: (2026)
Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
von: Zhang, David Junhao, et al.
Veröffentlicht: (2023)
von: Zhang, David Junhao, et al.
Veröffentlicht: (2023)
HairWeaver: Few-Shot Photorealistic Hair Motion Synthesis with Sim-to-Real Guided Video Diffusion
von: Chang, Di, et al.
Veröffentlicht: (2026)
von: Chang, Di, et al.
Veröffentlicht: (2026)
ChopGrad: Pixel-Wise Losses for Latent Video Diffusion via Truncated Backpropagation
von: Rivkin, Dmitriy, et al.
Veröffentlicht: (2026)
von: Rivkin, Dmitriy, et al.
Veröffentlicht: (2026)
Filtering Memorization from Parameter-Space in Diffusion Models
von: Zhe, Yu, et al.
Veröffentlicht: (2026)
von: Zhe, Yu, et al.
Veröffentlicht: (2026)
PixelGen: Improving Pixel Diffusion with Perceptual Supervision
von: Ma, Zehong, et al.
Veröffentlicht: (2026)
von: Ma, Zehong, et al.
Veröffentlicht: (2026)
Latent Watermark: Inject and Detect Watermarks in Latent Diffusion Space
von: Meng, Zheling, et al.
Veröffentlicht: (2024)
von: Meng, Zheling, et al.
Veröffentlicht: (2024)
Imagine Flash: Accelerating Emu Diffusion Models with Backward Distillation
von: Kohler, Jonas, et al.
Veröffentlicht: (2024)
von: Kohler, Jonas, et al.
Veröffentlicht: (2024)
One Small Step in Latent, One Giant Leap for Pixels: Fast Latent Upscale Adapter for Your Diffusion Models
von: Razin, Aleksandr, et al.
Veröffentlicht: (2025)
von: Razin, Aleksandr, et al.
Veröffentlicht: (2025)
StreamDiT: Real-Time Streaming Text-to-Video Generation
von: Kodaira, Akio, et al.
Veröffentlicht: (2025)
von: Kodaira, Akio, et al.
Veröffentlicht: (2025)
Improving Chain-of-Thought Efficiency for Autoregressive Image Generation
von: Gu, Zeqi, et al.
Veröffentlicht: (2025)
von: Gu, Zeqi, et al.
Veröffentlicht: (2025)
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
von: Cai, Yuanhao, et al.
Veröffentlicht: (2025)
von: Cai, Yuanhao, et al.
Veröffentlicht: (2025)
AVID: Any-Length Video Inpainting with Diffusion Model
von: Zhang, Zhixing, et al.
Veröffentlicht: (2023)
von: Zhang, Zhixing, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Cache Me if You Can: Accelerating Diffusion Models through Block Caching
von: Wimbauer, Felix, et al.
Veröffentlicht: (2023) -
Llama Learns to Direct: DirectorLLM for Human-Centric Video Generation
von: Song, Kunpeng, et al.
Veröffentlicht: (2024) -
Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
von: Ma, Xu, et al.
Veröffentlicht: (2025) -
DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space
von: He, Wenkun, et al.
Veröffentlicht: (2025) -
MoCha: Towards Movie-Grade Talking Character Synthesis
von: Wei, Cong, et al.
Veröffentlicht: (2025)