DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Issachar, Noam, Yariv, Guy, Benaim, Sagie, Adi, Yossi, Lischinski, Dani, Fattal, Raanan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915760275718144
author Issachar, Noam
Yariv, Guy
Benaim, Sagie
Adi, Yossi
Lischinski, Dani
Fattal, Raanan
author_facet Issachar, Noam
Yariv, Guy
Benaim, Sagie
Adi, Yossi
Lischinski, Dani
Fattal, Raanan
contents Diffusion Transformer models can generate images with remarkable fidelity and detail, yet training them at ultra-high resolutions remains extremely costly due to the self-attention mechanism's quadratic scaling with the number of image tokens. In this paper, we introduce Dynamic Position Extrapolation (DyPE), a novel, training-free method that enables pre-trained diffusion transformers to synthesize images at resolutions far beyond their training data, with no additional sampling cost. DyPE takes advantage of the spectral progression inherent to the diffusion process, where low-frequency structures converge early, while high-frequencies take more steps to resolve. Specifically, DyPE dynamically adjusts the model's positional encoding at each diffusion step, matching their frequency spectrum with the current stage of the generative process. This approach allows us to generate images at resolutions that exceed the training resolution dramatically, e.g., 16 million pixels using FLUX. On multiple benchmarks, DyPE consistently improves performance and achieves state-of-the-art fidelity in ultra-high-resolution image generation, with gains becoming even more pronounced at higher resolutions. Project page is available at https://noamissachar.github.io/DyPE/.
format Preprint
id arxiv_https___arxiv_org_abs_2510_20766
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion
Issachar, Noam
Yariv, Guy
Benaim, Sagie
Adi, Yossi
Lischinski, Dani
Fattal, Raanan
Computer Vision and Pattern Recognition
Diffusion Transformer models can generate images with remarkable fidelity and detail, yet training them at ultra-high resolutions remains extremely costly due to the self-attention mechanism's quadratic scaling with the number of image tokens. In this paper, we introduce Dynamic Position Extrapolation (DyPE), a novel, training-free method that enables pre-trained diffusion transformers to synthesize images at resolutions far beyond their training data, with no additional sampling cost. DyPE takes advantage of the spectral progression inherent to the diffusion process, where low-frequency structures converge early, while high-frequencies take more steps to resolve. Specifically, DyPE dynamically adjusts the model's positional encoding at each diffusion step, matching their frequency spectrum with the current stage of the generative process. This approach allows us to generate images at resolutions that exceed the training resolution dramatically, e.g., 16 million pixels using FLUX. On multiple benchmarks, DyPE consistently improves performance and achieves state-of-the-art fidelity in ultra-high-resolution image generation, with gains becoming even more pronounced at higher resolutions. Project page is available at https://noamissachar.github.io/DyPE/.
title DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.20766