More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918173393027072 |
|---|---|
| author | Lin, Hongkai Liang, Dingkang Du, Mingyang Zhou, Xin Bai, Xiang |
| author_facet | Lin, Hongkai Liang, Dingkang Du, Mingyang Zhou, Xin Bai, Xiang |
| contents | Generative depth estimation methods leverage the rich visual priors stored in pre-trained text-to-image diffusion models, demonstrating astonishing zero-shot capability. However, parameter updates during training lead to catastrophic degradation in the image generation capability of the pre-trained model. We introduce MERGE, a unified model for image generation and depth estimation, starting from a fixed pre-trained text-to-image model. MERGE demonstrates that the pre-trained text-to-image model can do more than image generation, but also expand to depth estimation effortlessly. Specifically, MERGE introduces a play-and-plug framework that enables seamless switching between image generation and depth estimation modes through simple and pluggable converters. Meanwhile, we propose a Group Reuse Mechanism to encourage parameter reuse and improve the utilization of the additional learnable parameters. MERGE unleashes the powerful depth estimation capability of the pre-trained text-to-image model while preserving its original image generation ability. Compared to other unified models for image generation and depth estimation, MERGE achieves state-of-the-art performance across multiple depth estimation benchmarks. The code will be made available at https://github.com/H-EmbodVis/MERGE |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_23574 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models Lin, Hongkai Liang, Dingkang Du, Mingyang Zhou, Xin Bai, Xiang Computer Vision and Pattern Recognition Generative depth estimation methods leverage the rich visual priors stored in pre-trained text-to-image diffusion models, demonstrating astonishing zero-shot capability. However, parameter updates during training lead to catastrophic degradation in the image generation capability of the pre-trained model. We introduce MERGE, a unified model for image generation and depth estimation, starting from a fixed pre-trained text-to-image model. MERGE demonstrates that the pre-trained text-to-image model can do more than image generation, but also expand to depth estimation effortlessly. Specifically, MERGE introduces a play-and-plug framework that enables seamless switching between image generation and depth estimation modes through simple and pluggable converters. Meanwhile, we propose a Group Reuse Mechanism to encourage parameter reuse and improve the utilization of the additional learnable parameters. MERGE unleashes the powerful depth estimation capability of the pre-trained text-to-image model while preserving its original image generation ability. Compared to other unified models for image generation and depth estimation, MERGE achieves state-of-the-art performance across multiple depth estimation benchmarks. The code will be made available at https://github.com/H-EmbodVis/MERGE |
| title | More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2510.23574 |