Exploring Diffusion Transformer Designs via Grafting
Fuente:
arXiv
Saved in:
| Main Authors: | Chandrasegaran, Keshigeyan, Poli, Michael, Fu, Daniel Y., Kim, Dongjun, Hadzic, Lea M., Li, Manling, Gupta, Agrim, Massaroli, Stefano, Mirhoseini, Azalia, Niebles, Juan Carlos, Ermon, Stefano, Fei-Fei, Li |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HourVideo: 1-Hour Video-Language Understanding
by: Chandrasegaran, Keshigeyan, et al.
Published: (2024)
by: Chandrasegaran, Keshigeyan, et al.
Published: (2024)
GPIC: A Giant Permissive Image Corpus for Visual Generation
by: Chandrasegaran, Keshigeyan, et al.
Published: (2026)
by: Chandrasegaran, Keshigeyan, et al.
Published: (2026)
T*: Re-thinking Temporal Search for Long-Form Video Understanding
by: Ye, Jinhui, et al.
Published: (2025)
by: Ye, Jinhui, et al.
Published: (2025)
Training-Free Safe Denoisers for Safe Use of Diffusion Models
by: Kim, Mingyu, et al.
Published: (2025)
by: Kim, Mingyu, et al.
Published: (2025)
That Chip Has Sailed: A Critique of Unfounded Skepticism Around AI for Chip Design
by: Goldie, Anna, et al.
Published: (2024)
by: Goldie, Anna, et al.
Published: (2024)
The Principles of Diffusion Models
by: Lai, Chieh-Hsin, et al.
Published: (2025)
by: Lai, Chieh-Hsin, et al.
Published: (2025)
On the Role of Temperature Sampling in Test-Time Scaling
by: Wu, Yuheng, et al.
Published: (2025)
by: Wu, Yuheng, et al.
Published: (2025)
Changen2: Multi-Temporal Remote Sensing Generative Change Foundation Model
by: Zheng, Zhuo, et al.
Published: (2024)
by: Zheng, Zhuo, et al.
Published: (2024)
Mechanistic Design and Scaling of Hybrid Architectures
by: Poli, Michael, et al.
Published: (2024)
by: Poli, Michael, et al.
Published: (2024)
Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
by: Zhang, Pingyue, et al.
Published: (2026)
by: Zhang, Pingyue, et al.
Published: (2026)
MindCube: Spatial Mental Modeling from Limited Views
by: Wang, Qineng, et al.
Published: (2025)
by: Wang, Qineng, et al.
Published: (2025)
STAR: Synthesis of Tailored Architectures
by: Thomas, Armin W., et al.
Published: (2024)
by: Thomas, Armin W., et al.
Published: (2024)
Think, Prune, Train, Improve: Scaling Reasoning without Scaling Models
by: Costello, Caia, et al.
Published: (2025)
by: Costello, Caia, et al.
Published: (2025)
Sliding Window Recurrences for Sequence Models
by: Secrieru, Dragos, et al.
Published: (2025)
by: Secrieru, Dragos, et al.
Published: (2025)
Hydragen: High-Throughput LLM Inference with Shared Prefixes
by: Juravsky, Jordan, et al.
Published: (2024)
by: Juravsky, Jordan, et al.
Published: (2024)
Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution
by: Lou, Aaron, et al.
Published: (2023)
by: Lou, Aaron, et al.
Published: (2023)
Federation of Experts: Communication Efficient Distributed Inference for Large Language Models
by: Abdurrahman, Muhammad Shahir, et al.
Published: (2026)
by: Abdurrahman, Muhammad Shahir, et al.
Published: (2026)
TRACE: Capability-Targeted Agentic Training
by: Kang, Hangoo, et al.
Published: (2026)
by: Kang, Hangoo, et al.
Published: (2026)
ForTIFAI: Fending Off Recursive Training Induced Failure for AI Model Collapse
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2025)
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2025)
Agent JIT Compilation for Latency-Optimizing Web Agent Planning and Scheduling
by: Winston, Caleb, et al.
Published: (2026)
by: Winston, Caleb, et al.
Published: (2026)
Unsupervised Anomaly Detection Using Diffusion Trend Analysis for Display Inspection
by: Kim, Eunwoo, et al.
Published: (2024)
by: Kim, Eunwoo, et al.
Published: (2024)
Reviving Any-Subset Autoregressive Models with Principled Parallel Sampling and Speculative Decoding
by: Guo, Gabe, et al.
Published: (2025)
by: Guo, Gabe, et al.
Published: (2025)
SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking
by: Cundy, Chris, et al.
Published: (2023)
by: Cundy, Chris, et al.
Published: (2023)
Model Inversion Robustness: Can Transfer Learning Help?
by: Ho, Sy-Tuyen, et al.
Published: (2024)
by: Ho, Sy-Tuyen, et al.
Published: (2024)
A Survey on Generative Modeling with Limited Data, Few Shots, and Zero Shot
by: Abdollahzadeh, Milad, et al.
Published: (2023)
by: Abdollahzadeh, Milad, et al.
Published: (2023)
Self-Refining Diffusion Samplers: Enabling Parallelization via Parareal Iterations
by: Selvam, Nikil Roashan, et al.
Published: (2024)
by: Selvam, Nikil Roashan, et al.
Published: (2024)
PaGoDA: Progressive Growing of a One-Step Generator from a Low-Resolution Diffusion Teacher
by: Kim, Dongjun, et al.
Published: (2024)
by: Kim, Dongjun, et al.
Published: (2024)
Data Unlearning in Diffusion Models
by: Alberti, Silas, et al.
Published: (2025)
by: Alberti, Silas, et al.
Published: (2025)
Divergence Minimization Preference Optimization for Diffusion Model Alignment
by: Li, Binxu, et al.
Published: (2025)
by: Li, Binxu, et al.
Published: (2025)
State-Free Inference of State-Space Models: The Transfer Function Approach
by: Parnichkun, Rom N., et al.
Published: (2024)
by: Parnichkun, Rom N., et al.
Published: (2024)
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Consistency Trajectory Models: Learning Probability Flow ODE Trajectory of Diffusion
by: Kim, Dongjun, et al.
Published: (2023)
by: Kim, Dongjun, et al.
Published: (2023)
Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models
by: Wang, Austin, et al.
Published: (2026)
by: Wang, Austin, et al.
Published: (2026)
Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use
by: Goldie, Anna, et al.
Published: (2025)
by: Goldie, Anna, et al.
Published: (2025)
CHESS: Contextual Harnessing for Efficient SQL Synthesis
by: Talaei, Shayan, et al.
Published: (2024)
by: Talaei, Shayan, et al.
Published: (2024)
CATS: Contextually-Aware Thresholding for Sparsity in Large Language Models
by: Lee, Donghyun, et al.
Published: (2024)
by: Lee, Donghyun, et al.
Published: (2024)
RFG: Test-Time Scaling for Diffusion Large Language Model Reasoning with Reward-Free Guidance
by: Chen, Tianlang, et al.
Published: (2025)
by: Chen, Tianlang, et al.
Published: (2025)
Uncovering Emotion in Youth Digital Civic Participation Around Climate Change: Entanglements of Fear, Despair, and Anger in Civic Practice
by: Lynne Zummo, et al.
Published: (2025)
by: Lynne Zummo, et al.
Published: (2025)
Geometric Trajectory Diffusion Models
by: Han, Jiaqi, et al.
Published: (2024)
by: Han, Jiaqi, et al.
Published: (2024)
Educação Permanente para o aperfeiçoamento do Controle de Infecção Hospitalar: revisão integrativa
by: Aline Massaroli
Published: (2014)
by: Aline Massaroli
Published: (2014)
Similar Items
-
HourVideo: 1-Hour Video-Language Understanding
by: Chandrasegaran, Keshigeyan, et al.
Published: (2024) -
GPIC: A Giant Permissive Image Corpus for Visual Generation
by: Chandrasegaran, Keshigeyan, et al.
Published: (2026) -
T*: Re-thinking Temporal Search for Long-Form Video Understanding
by: Ye, Jinhui, et al.
Published: (2025) -
Training-Free Safe Denoisers for Safe Use of Diffusion Models
by: Kim, Mingyu, et al.
Published: (2025) -
That Chip Has Sailed: A Critique of Unfounded Skepticism Around AI for Chip Design
by: Goldie, Anna, et al.
Published: (2024)