MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model
Fuente:
arXiv
Saved in:
| Main Authors: | Dao, Quan, Metaxas, Dimitris |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Text-to-Image Generation
by: Dao, Quan, et al.
Published: (2024)
by: Dao, Quan, et al.
Published: (2024)
DiMSUM: Diffusion Mamba -- A Scalable and Unified Spatial-Frequency Method for Image Generation
by: Phung, Hao, et al.
Published: (2024)
by: Phung, Hao, et al.
Published: (2024)
Improved Training Technique for Latent Consistency Models
by: Dao, Quan, et al.
Published: (2025)
by: Dao, Quan, et al.
Published: (2025)
AutoEdit: Automatic Hyperparameter Tuning for Image Editing
by: Pham, Chau, et al.
Published: (2025)
by: Pham, Chau, et al.
Published: (2025)
Score-Guided Diffusion for 3D Human Recovery
by: Stathopoulos, Anastasis, et al.
Published: (2024)
by: Stathopoulos, Anastasis, et al.
Published: (2024)
Overcoming the Curvature Bottleneck in MeanFlow
by: Zhang, Xinxi, et al.
Published: (2025)
by: Zhang, Xinxi, et al.
Published: (2025)
StreamFlow: Theory, Algorithm, and Implementation for High-Efficiency Rectified Flow Generation
by: Fang, Sen, et al.
Published: (2025)
by: Fang, Sen, et al.
Published: (2025)
Steering Rectified Flow Models in the Vector Field for Controlled Image Generation
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
RAC: Rectified Flow Auto Coder
by: Fang, Sen, et al.
Published: (2026)
by: Fang, Sen, et al.
Published: (2026)
Improving Visual Reasoning with Iterative Evidence Refinement
by: Shi, Zeru, et al.
Published: (2026)
by: Shi, Zeru, et al.
Published: (2026)
SINE: SINgle Image Editing with Text-to-Image Diffusion Models
by: Zhang, Zhixing, et al.
Published: (2022)
by: Zhang, Zhixing, et al.
Published: (2022)
DeDPO: Debiased Direct Preference Optimization for Diffusion Models
by: Pham, Khiem, et al.
Published: (2026)
by: Pham, Khiem, et al.
Published: (2026)
Resolving Inconsistent Semantics in Multi-Dataset Image Segmentation
by: Zhangli, Qilong, et al.
Published: (2024)
by: Zhangli, Qilong, et al.
Published: (2024)
Learning to Localize Actions in Instructional Videos with LLM-Based Multi-Pathway Text-Video Alignment
by: Chen, Yuxiao, et al.
Published: (2024)
by: Chen, Yuxiao, et al.
Published: (2024)
FlowIt: Global Matching via Hierarchical Transformers and Optimal Transport for Optical Flow
by: Safadoust, Sadra, et al.
Published: (2026)
by: Safadoust, Sadra, et al.
Published: (2026)
Dynamic Texture Transfer using PatchMatch and Transformers
by: Pu, Guo, et al.
Published: (2024)
by: Pu, Guo, et al.
Published: (2024)
PatchScaler: An Efficient Patch-Independent Diffusion Model for Image Super-Resolution
by: Liu, Yong, et al.
Published: (2024)
by: Liu, Yong, et al.
Published: (2024)
Neural Deformable Models for 3D Bi-Ventricular Heart Shape Reconstruction and Modeling from 2D Sparse Cardiac Magnetic Resonance Imaging
by: Ye, Meng, et al.
Published: (2023)
by: Ye, Meng, et al.
Published: (2023)
Continuous Spatio-Temporal Memory Networks for 4D Cardiac Cine MRI Segmentation
by: Ye, Meng, et al.
Published: (2024)
by: Ye, Meng, et al.
Published: (2024)
Spectrum-Aware Parameter Efficient Fine-Tuning for Diffusion Models
by: Zhang, Xinxi, et al.
Published: (2024)
by: Zhang, Xinxi, et al.
Published: (2024)
AVID: Any-Length Video Inpainting with Diffusion Model
by: Zhang, Zhixing, et al.
Published: (2023)
by: Zhang, Zhixing, et al.
Published: (2023)
Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
by: Dao, Quan, et al.
Published: (2025)
by: Dao, Quan, et al.
Published: (2025)
DDiT: Dynamic Patch Scheduling for Efficient Diffusion Transformers
by: Kim, Dahye, et al.
Published: (2026)
by: Kim, Dahye, et al.
Published: (2026)
PATS: Patch Area Transportation with Subdivision for Local Feature Matching
by: Ni, Junjie, et al.
Published: (2023)
by: Ni, Junjie, et al.
Published: (2023)
Test-Time Spectrum-Aware Latent Steering for Zero-Shot Generalization in Vision-Language Models
by: Dafnis, Konstantinos M., et al.
Published: (2025)
by: Dafnis, Konstantinos M., et al.
Published: (2025)
An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental Learning
by: Tran, Quyen, et al.
Published: (2022)
by: Tran, Quyen, et al.
Published: (2022)
Layout Agnostic Scene Text Image Synthesis with Diffusion Models
by: Zhangli, Qilong, et al.
Published: (2024)
by: Zhangli, Qilong, et al.
Published: (2024)
Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality Generation
by: Park, Dogyun, et al.
Published: (2025)
by: Park, Dogyun, et al.
Published: (2025)
Instantaneous Perception of Moving Objects in 3D
by: Liu, Di, et al.
Published: (2024)
by: Liu, Di, et al.
Published: (2024)
Relational Representation Learning Network for Cross-Spectral Image Patch Matching
by: Yu, Chuang, et al.
Published: (2024)
by: Yu, Chuang, et al.
Published: (2024)
Local Patches Meet Global Context: Scalable 3D Diffusion Priors for Computed Tomography Reconstruction
by: Yang, Taewon, et al.
Published: (2025)
by: Yang, Taewon, et al.
Published: (2025)
DIAGNOSIS: Detecting Unauthorized Data Usages in Text-to-image Diffusion Models
by: Wang, Zhenting, et al.
Published: (2023)
by: Wang, Zhenting, et al.
Published: (2023)
An Optimized PatchMatch for Multi-scale and Multi-feature Label Fusion
by: Giraud, Rémi, et al.
Published: (2019)
by: Giraud, Rémi, et al.
Published: (2019)
Beyond Pixels: Semi-Supervised Semantic Segmentation with a Multi-scale Patch-based Multi-Label Classifier
by: Howlader, Prantik, et al.
Published: (2024)
by: Howlader, Prantik, et al.
Published: (2024)
ERUPT: Efficient Rendering with Unposed Patch Transformer
by: Shugaev, Maxim V., et al.
Published: (2025)
by: Shugaev, Maxim V., et al.
Published: (2025)
Learning Volumetric Neural Deformable Models to Recover 3D Regional Heart Wall Motion from Multi-Planar Tagged MRI
by: Ye, Meng, et al.
Published: (2024)
by: Ye, Meng, et al.
Published: (2024)
A High-Quality Robust Diffusion Framework for Corrupted Dataset
by: Dao, Quan, et al.
Published: (2023)
by: Dao, Quan, et al.
Published: (2023)
Boosting Latent Diffusion with Flow Matching
by: Schusterbauer, Johannes, et al.
Published: (2023)
by: Schusterbauer, Johannes, et al.
Published: (2023)
LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation
by: Jin, Can, et al.
Published: (2025)
by: Jin, Can, et al.
Published: (2025)
Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation
by: Zou, Xuechao, et al.
Published: (2025)
by: Zou, Xuechao, et al.
Published: (2025)
Similar Items
-
Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Text-to-Image Generation
by: Dao, Quan, et al.
Published: (2024) -
DiMSUM: Diffusion Mamba -- A Scalable and Unified Spatial-Frequency Method for Image Generation
by: Phung, Hao, et al.
Published: (2024) -
Improved Training Technique for Latent Consistency Models
by: Dao, Quan, et al.
Published: (2025) -
AutoEdit: Automatic Hyperparameter Tuning for Image Editing
by: Pham, Chau, et al.
Published: (2025) -
Score-Guided Diffusion for 3D Human Recovery
by: Stathopoulos, Anastasis, et al.
Published: (2024)