Semantic and Temporal Integration in Latent Diffusion Space for High-Fidelity Video Super-Resolution
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915421732470784 |
|---|---|
| author | Wang, Yiwen Chai, Xinning Zhang, Yuhong Cheng, Zhengxue Zhao, Jun Xie, Rong Song, Li |
| author_facet | Wang, Yiwen Chai, Xinning Zhang, Yuhong Cheng, Zhengxue Zhao, Jun Xie, Rong Song, Li |
| contents | Recent advancements in video super-resolution (VSR) models have demonstrated impressive results in enhancing low-resolution videos. However, due to limitations in adequately controlling the generation process, achieving high fidelity alignment with the low-resolution input while maintaining temporal consistency across frames remains a significant challenge. In this work, we propose Semantic and Temporal Guided Video Super-Resolution (SeTe-VSR), a novel approach that incorporates both semantic and temporal-spatio guidance in the latent diffusion space to address these challenges. By incorporating high-level semantic information and integrating spatial and temporal information, our approach achieves a seamless balance between recovering intricate details and ensuring temporal coherence. Our method not only preserves high-reality visual content but also significantly enhances fidelity. Extensive experiments demonstrate that SeTe-VSR outperforms existing methods in terms of detail recovery and perceptual quality, highlighting its effectiveness for complex video super-resolution tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_00471 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Semantic and Temporal Integration in Latent Diffusion Space for High-Fidelity Video Super-Resolution Wang, Yiwen Chai, Xinning Zhang, Yuhong Cheng, Zhengxue Zhao, Jun Xie, Rong Song, Li Computer Vision and Pattern Recognition Image and Video Processing Recent advancements in video super-resolution (VSR) models have demonstrated impressive results in enhancing low-resolution videos. However, due to limitations in adequately controlling the generation process, achieving high fidelity alignment with the low-resolution input while maintaining temporal consistency across frames remains a significant challenge. In this work, we propose Semantic and Temporal Guided Video Super-Resolution (SeTe-VSR), a novel approach that incorporates both semantic and temporal-spatio guidance in the latent diffusion space to address these challenges. By incorporating high-level semantic information and integrating spatial and temporal information, our approach achieves a seamless balance between recovering intricate details and ensuring temporal coherence. Our method not only preserves high-reality visual content but also significantly enhances fidelity. Extensive experiments demonstrate that SeTe-VSR outperforms existing methods in terms of detail recovery and perceptual quality, highlighting its effectiveness for complex video super-resolution tasks. |
| title | Semantic and Temporal Integration in Latent Diffusion Space for High-Fidelity Video Super-Resolution |
| topic | Computer Vision and Pattern Recognition Image and Video Processing |
| url | https://arxiv.org/abs/2508.00471 |