On Inductive Biases That Enable Generalization of Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | An, Jie, Wang, De, Guo, Pengsheng, Luo, Jiebo, Schwing, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Variational Rectified Flow Matching
by: Guo, Pengsheng, et al.
Published: (2025)
by: Guo, Pengsheng, et al.
Published: (2025)
Bridging Diffusion Models and 3D Representations: A 3D Consistent Super-Resolution Framework
by: Chen, Yi-Ting, et al.
Published: (2025)
by: Chen, Yi-Ting, et al.
Published: (2025)
Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training
by: Zhou, Zhenghong, et al.
Published: (2024)
by: Zhou, Zhenghong, et al.
Published: (2024)
PixelDiT: Pixel Diffusion Transformers for Image Generation
by: Yu, Yongsheng, et al.
Published: (2025)
by: Yu, Yongsheng, et al.
Published: (2025)
Diffusion Transformer-to-Mamba Distillation for High-Resolution Image Generation
by: Yao, Yuan, et al.
Published: (2025)
by: Yao, Yuan, et al.
Published: (2025)
Ouroboros-Diffusion: Exploring Consistent Content Generation in Tuning-free Long Video Diffusion
by: Chen, Jingyuan, et al.
Published: (2025)
by: Chen, Jingyuan, et al.
Published: (2025)
Confronting Reward Overoptimization for Diffusion Models: A Perspective of Inductive and Primacy Biases
by: Zhang, Ziyi, et al.
Published: (2024)
by: Zhang, Ziyi, et al.
Published: (2024)
Rethinking Inductive Biases for Surface Normal Estimation
by: Bae, Gwangbin, et al.
Published: (2024)
by: Bae, Gwangbin, et al.
Published: (2024)
Bring Metric Functions into Diffusion Models
by: An, Jie, et al.
Published: (2024)
by: An, Jie, et al.
Published: (2024)
VideoSketcher: Video Models Prior Enable Versatile Sequential Sketch Generation
by: Ren, Hui, et al.
Published: (2026)
by: Ren, Hui, et al.
Published: (2026)
The Curse of Conditions: Analyzing and Improving Optimal Transport for Conditional Flow-Based Generation
by: Cheng, Ho Kei, et al.
Published: (2025)
by: Cheng, Ho Kei, et al.
Published: (2025)
CAR-Flow: Condition-Aware Reparameterization Aligns Source and Target for Better Flow Matching
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
NoPo-Avatar: Generalizable and Animatable Avatars from Sparse Inputs without Human Poses
by: Wen, Jing, et al.
Published: (2025)
by: Wen, Jing, et al.
Published: (2025)
LIFe-GoM: Generalizable Human Rendering with Learned Iterative Feedback Over Multi-Resolution Gaussians-on-Mesh
by: Wen, Jing, et al.
Published: (2025)
by: Wen, Jing, et al.
Published: (2025)
SimpliHuMoN: Simplifying Human Motion Prediction
by: Agrawal, Aadya, et al.
Published: (2026)
by: Agrawal, Aadya, et al.
Published: (2026)
Leveraging Geometric Visual Illusions as Perceptual Inductive Biases for Vision Models
by: Yang, Haobo, et al.
Published: (2025)
by: Yang, Haobo, et al.
Published: (2025)
RING-NeRF : Rethinking Inductive Biases for Versatile and Efficient Neural Fields
by: Petit, Doriand, et al.
Published: (2023)
by: Petit, Doriand, et al.
Published: (2023)
ZipIR: Latent Pyramid Diffusion Transformer for High-Resolution Image Restoration
by: Yu, Yongsheng, et al.
Published: (2025)
by: Yu, Yongsheng, et al.
Published: (2025)
RELOCATE: A Simple Training-Free Baseline for Visual Query Localization Using Region-Based Representations
by: Khosla, Savya, et al.
Published: (2024)
by: Khosla, Savya, et al.
Published: (2024)
Tripod: Three Complementary Inductive Biases for Disentangled Representation Learning
by: Hsu, Kyle, et al.
Published: (2024)
by: Hsu, Kyle, et al.
Published: (2024)
MIRA: Multimodal Iterative Reasoning Agent for Image Editing
by: Zeng, Ziyun, et al.
Published: (2025)
by: Zeng, Ziyun, et al.
Published: (2025)
Continuous-Multiple Image Outpainting in One-Step via Positional Query and A Diffusion-based Approach
by: Zhang, Shaofeng, et al.
Published: (2024)
by: Zhang, Shaofeng, et al.
Published: (2024)
Chain-of-Thought Prompting for Demographic Inference with Large Multimodal Models
by: Yu, Yongsheng, et al.
Published: (2024)
by: Yu, Yongsheng, et al.
Published: (2024)
A Versatile Multimodal Agent for Multimedia Content Generation
by: Zhang, Daoan, et al.
Published: (2026)
by: Zhang, Daoan, et al.
Published: (2026)
Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
by: Liu, Shizhan, et al.
Published: (2025)
by: Liu, Shizhan, et al.
Published: (2025)
End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting
by: Wang, Yongqi, et al.
Published: (2024)
by: Wang, Yongqi, et al.
Published: (2024)
Simulate, Refocus and Ensemble: An Attention-Refocusing Scheme for Domain Generalization
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement
by: Zheng, Haitian, et al.
Published: (2025)
by: Zheng, Haitian, et al.
Published: (2025)
Holistic Visual-Textual Sentiment Analysis with Prior Models
by: Chen, Junyu, et al.
Published: (2022)
by: Chen, Junyu, et al.
Published: (2022)
Object-Centric Temporal Consistency via Conditional Autoregressive Inductive Biases
by: Meo, Cristian, et al.
Published: (2024)
by: Meo, Cristian, et al.
Published: (2024)
GoMAvatar: Efficient Animatable Human Modeling from Monocular Video Using Gaussians-on-Mesh
by: Wen, Jing, et al.
Published: (2024)
by: Wen, Jing, et al.
Published: (2024)
VFX Creator: Animated Visual Effect Generation with Controllable Diffusion Transformer
by: Liu, Xinyu, et al.
Published: (2025)
by: Liu, Xinyu, et al.
Published: (2025)
Z-Erase: Enabling Concept Erasure in Single-Stream Diffusion Transformers
by: Jiang, Nanxiang, et al.
Published: (2026)
by: Jiang, Nanxiang, et al.
Published: (2026)
Elastic Diffusion Transformer
by: Wang, Jiangshan, et al.
Published: (2026)
by: Wang, Jiangshan, et al.
Published: (2026)
Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report Generation
by: Chen, Wenting, et al.
Published: (2023)
by: Chen, Wenting, et al.
Published: (2023)
Studying Classifier(-Free) Guidance From a Classifier-Centric Perspective
by: Zhao, Xiaoming, et al.
Published: (2025)
by: Zhao, Xiaoming, et al.
Published: (2025)
Low-Biased General Annotated Dataset Generation
by: Jiang, Dengyang, et al.
Published: (2024)
by: Jiang, Dengyang, et al.
Published: (2024)
OW-VISCapTor: Abstractors for Open-World Video Instance Segmentation and Captioning
by: Choudhuri, Anwesa, et al.
Published: (2024)
by: Choudhuri, Anwesa, et al.
Published: (2024)
REN: Fast and Efficient Region Encodings from Patch-Based Image Encoders
by: Khosla, Savya, et al.
Published: (2025)
by: Khosla, Savya, et al.
Published: (2025)
HOIGen-1M: A Large-scale Dataset for Human-Object Interaction Video Generation
by: Liu, Kun, et al.
Published: (2025)
by: Liu, Kun, et al.
Published: (2025)
Similar Items
-
Variational Rectified Flow Matching
by: Guo, Pengsheng, et al.
Published: (2025) -
Bridging Diffusion Models and 3D Representations: A 3D Consistent Super-Resolution Framework
by: Chen, Yi-Ting, et al.
Published: (2025) -
Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training
by: Zhou, Zhenghong, et al.
Published: (2024) -
PixelDiT: Pixel Diffusion Transformers for Image Generation
by: Yu, Yongsheng, et al.
Published: (2025) -
Diffusion Transformer-to-Mamba Distillation for High-Resolution Image Generation
by: Yao, Yuan, et al.
Published: (2025)