Semantic and Temporal Integration in Latent Diffusion Space for High-Fidelity Video Super-Resolution

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Yiwen, Chai, Xinning, Zhang, Yuhong, Cheng, Zhengxue, Zhao, Jun, Xie, Rong, Song, Li
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915421732470784
author Wang, Yiwen
Chai, Xinning
Zhang, Yuhong
Cheng, Zhengxue
Zhao, Jun
Xie, Rong
Song, Li
author_facet Wang, Yiwen
Chai, Xinning
Zhang, Yuhong
Cheng, Zhengxue
Zhao, Jun
Xie, Rong
Song, Li
contents Recent advancements in video super-resolution (VSR) models have demonstrated impressive results in enhancing low-resolution videos. However, due to limitations in adequately controlling the generation process, achieving high fidelity alignment with the low-resolution input while maintaining temporal consistency across frames remains a significant challenge. In this work, we propose Semantic and Temporal Guided Video Super-Resolution (SeTe-VSR), a novel approach that incorporates both semantic and temporal-spatio guidance in the latent diffusion space to address these challenges. By incorporating high-level semantic information and integrating spatial and temporal information, our approach achieves a seamless balance between recovering intricate details and ensuring temporal coherence. Our method not only preserves high-reality visual content but also significantly enhances fidelity. Extensive experiments demonstrate that SeTe-VSR outperforms existing methods in terms of detail recovery and perceptual quality, highlighting its effectiveness for complex video super-resolution tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00471
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Semantic and Temporal Integration in Latent Diffusion Space for High-Fidelity Video Super-Resolution
Wang, Yiwen
Chai, Xinning
Zhang, Yuhong
Cheng, Zhengxue
Zhao, Jun
Xie, Rong
Song, Li
Computer Vision and Pattern Recognition
Image and Video Processing
Recent advancements in video super-resolution (VSR) models have demonstrated impressive results in enhancing low-resolution videos. However, due to limitations in adequately controlling the generation process, achieving high fidelity alignment with the low-resolution input while maintaining temporal consistency across frames remains a significant challenge. In this work, we propose Semantic and Temporal Guided Video Super-Resolution (SeTe-VSR), a novel approach that incorporates both semantic and temporal-spatio guidance in the latent diffusion space to address these challenges. By incorporating high-level semantic information and integrating spatial and temporal information, our approach achieves a seamless balance between recovering intricate details and ensuring temporal coherence. Our method not only preserves high-reality visual content but also significantly enhances fidelity. Extensive experiments demonstrate that SeTe-VSR outperforms existing methods in terms of detail recovery and perceptual quality, highlighting its effectiveness for complex video super-resolution tasks.
title Semantic and Temporal Integration in Latent Diffusion Space for High-Fidelity Video Super-Resolution
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2508.00471