A Simple Approach to Unifying Diffusion-based Conditional Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Xirui, Herrmann, Charles, Chan, Kelvin C. K., Li, Yinxiao, Sun, Deqing, Ma, Chao, Yang, Ming-Hsuan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
por: Zhang, Junyi, et al.
Publicado: (2024)
por: Zhang, Junyi, et al.
Publicado: (2024)
DynamicScaler: Seamless and Scalable Video Generation for Panoramic Scenes
por: Liu, Jinxiu, et al.
Publicado: (2024)
por: Liu, Jinxiu, et al.
Publicado: (2024)
HoliSDiP: Image Super-Resolution via Holistic Semantics and Diffusion Prior
por: Tsao, Li-Yuan, et al.
Publicado: (2024)
por: Tsao, Li-Yuan, et al.
Publicado: (2024)
CoCoIns: Consistent Subject Generation via Contrastive Instantiated Concepts
por: Hsin-Ying, Lee, et al.
Publicado: (2025)
por: Hsin-Ying, Lee, et al.
Publicado: (2025)
Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence
por: Zhang, Junyi, et al.
Publicado: (2023)
por: Zhang, Junyi, et al.
Publicado: (2023)
HoliGS: Holistic Gaussian Splatting for Embodied View Synthesis
por: Wang, Xiaoyuan, et al.
Publicado: (2025)
por: Wang, Xiaoyuan, et al.
Publicado: (2025)
From Prompt to Progression: Taming Video Diffusion Models for Seamless Attribute Transition
por: Lo, Ling, et al.
Publicado: (2025)
por: Lo, Ling, et al.
Publicado: (2025)
MASIV: Toward Material-Agnostic System Identification from Videos
por: Zhao, Yizhou, et al.
Publicado: (2025)
por: Zhao, Yizhou, et al.
Publicado: (2025)
Unified Dense Prediction of Video Diffusion
por: Yang, Lehan, et al.
Publicado: (2025)
por: Yang, Lehan, et al.
Publicado: (2025)
Calibrated Multi-Preference Optimization for Aligning Diffusion Models
por: Lee, Kyungmin, et al.
Publicado: (2025)
por: Lee, Kyungmin, et al.
Publicado: (2025)
Latent Knowledge-Guided Video Diffusion for Scientific Phenomena Generation from a Single Initial Frame
por: Cao, Qinglong, et al.
Publicado: (2024)
por: Cao, Qinglong, et al.
Publicado: (2024)
LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
por: Zhang, Junyi, et al.
Publicado: (2026)
por: Zhang, Junyi, et al.
Publicado: (2026)
Force Prompting: Video Generation Models Can Learn and Generalize Physics-based Control Signals
por: Gillman, Nate, et al.
Publicado: (2025)
por: Gillman, Nate, et al.
Publicado: (2025)
Improving Subject-Driven Image Synthesis with Subject-Agnostic Guidance
por: Chan, Kelvin C. K., et al.
Publicado: (2024)
por: Chan, Kelvin C. K., et al.
Publicado: (2024)
Reference Twice: A Simple and Unified Baseline for Few-Shot Instance Segmentation
por: Han, Yue, et al.
Publicado: (2023)
por: Han, Yue, et al.
Publicado: (2023)
High-Resolution Frame Interpolation with Patch-based Cascaded Diffusion
por: Hur, Junhwa, et al.
Publicado: (2024)
por: Hur, Junhwa, et al.
Publicado: (2024)
DrivingGaussian++: Towards Realistic Reconstruction and Editable Simulation for Surrounding Dynamic Driving Scenes
por: Xiong, Yajiao, et al.
Publicado: (2025)
por: Xiong, Yajiao, et al.
Publicado: (2025)
Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing
por: Yang, Hao, et al.
Publicado: (2026)
por: Yang, Hao, et al.
Publicado: (2026)
GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure
por: Gu, Leslie, et al.
Publicado: (2025)
por: Gu, Leslie, et al.
Publicado: (2025)
Multi-task Image Restoration Guided By Robust DINO Features
por: Lin, Xin, et al.
Publicado: (2023)
por: Lin, Xin, et al.
Publicado: (2023)
Effective Adapter for Face Recognition in the Wild
por: Liu, Yunhao, et al.
Publicado: (2023)
por: Liu, Yunhao, et al.
Publicado: (2023)
SERUM: Simple, Efficient, Robust, and Unifying Marking for Diffusion-based Image Generation
por: Kociszewski, Jan, et al.
Publicado: (2026)
por: Kociszewski, Jan, et al.
Publicado: (2026)
Pyramid Diffusion for Fine 3D Large Scene Generation
por: Liu, Yuheng, et al.
Publicado: (2023)
por: Liu, Yuheng, et al.
Publicado: (2023)
UFO-4D: Unposed Feedforward 4D Reconstruction from Two Images
por: Hur, Junhwa, et al.
Publicado: (2026)
por: Hur, Junhwa, et al.
Publicado: (2026)
GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting
por: Zhou, Xiaoyu, et al.
Publicado: (2024)
por: Zhou, Xiaoyu, et al.
Publicado: (2024)
UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion
por: Li, Wei, et al.
Publicado: (2024)
por: Li, Wei, et al.
Publicado: (2024)
EA3D: Online Open-World 3D Object Extraction from Streaming Videos
por: Zhou, Xiaoyu, et al.
Publicado: (2025)
por: Zhou, Xiaoyu, et al.
Publicado: (2025)
DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Autonomous Driving Scenes
por: Zhou, Xiaoyu, et al.
Publicado: (2023)
por: Zhou, Xiaoyu, et al.
Publicado: (2023)
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
por: Lin, Yuanze, et al.
Publicado: (2025)
por: Lin, Yuanze, et al.
Publicado: (2025)
Dual Associated Encoder for Face Restoration
por: Tsai, Yu-Ju, et al.
Publicado: (2023)
por: Tsai, Yu-Ju, et al.
Publicado: (2023)
Tex4D: Zero-shot 4D Scene Texturing with Video Diffusion Models
por: Bao, Jingzhi, et al.
Publicado: (2024)
por: Bao, Jingzhi, et al.
Publicado: (2024)
What does RL improve for Visual Reasoning? A Frankenstein-Style Analysis
por: Li, Xirui, et al.
Publicado: (2026)
por: Li, Xirui, et al.
Publicado: (2026)
Boundary Attention: Learning curves, corners, junctions and grouping
por: Polansky, Mia Gaia, et al.
Publicado: (2024)
por: Polansky, Mia Gaia, et al.
Publicado: (2024)
Tuning-Free Image Editing with Fidelity and Editability via Unified Latent Diffusion Model
por: Mao, Qi, et al.
Publicado: (2025)
por: Mao, Qi, et al.
Publicado: (2025)
Video Prediction Transformers without Recurrence or Convolution
por: Tang, Yujin, et al.
Publicado: (2024)
por: Tang, Yujin, et al.
Publicado: (2024)
Lumiere: A Space-Time Diffusion Model for Video Generation
por: Bar-Tal, Omer, et al.
Publicado: (2024)
por: Bar-Tal, Omer, et al.
Publicado: (2024)
Rethinking Video Tokenization: A Conditioned Diffusion-based Approach
por: Yang, Nianzu, et al.
Publicado: (2025)
por: Yang, Nianzu, et al.
Publicado: (2025)
UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis
por: Wang, Yuanrui, et al.
Publicado: (2025)
por: Wang, Yuanrui, et al.
Publicado: (2025)
UGD-IML: A Unified Generative Diffusion-based Framework for Constrained and Unconstrained Image Manipulation Localization
por: Mi, Yachun, et al.
Publicado: (2025)
por: Mi, Yachun, et al.
Publicado: (2025)
Re-boosting Self-Collaboration Parallel Prompt GAN for Unsupervised Image Restoration
por: Lin, Xin, et al.
Publicado: (2024)
por: Lin, Xin, et al.
Publicado: (2024)
Ejemplares similares
-
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
por: Zhang, Junyi, et al.
Publicado: (2024) -
DynamicScaler: Seamless and Scalable Video Generation for Panoramic Scenes
por: Liu, Jinxiu, et al.
Publicado: (2024) -
HoliSDiP: Image Super-Resolution via Holistic Semantics and Diffusion Prior
por: Tsao, Li-Yuan, et al.
Publicado: (2024) -
CoCoIns: Consistent Subject Generation via Contrastive Instantiated Concepts
por: Hsin-Ying, Lee, et al.
Publicado: (2025) -
Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence
por: Zhang, Junyi, et al.
Publicado: (2023)