Super-Resolution Generative Adversarial Networks based Video Enhancement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Çetin, Kağan, Akça, Hacer, Gerek, Ömer Nezih
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911027002605568
author Çetin, Kağan
Akça, Hacer
Gerek, Ömer Nezih
author_facet Çetin, Kağan
Akça, Hacer
Gerek, Ömer Nezih
contents This study introduces an enhanced approach to video super-resolution by extending ordinary Single-Image Super-Resolution (SISR) Super-Resolution Generative Adversarial Network (SRGAN) structure to handle spatio-temporal data. While SRGAN has proven effective for single-image enhancement, its design does not account for the temporal continuity required in video processing. To address this, a modified framework that incorporates 3D Non-Local Blocks is proposed, which is enabling the model to capture relationships across both spatial and temporal dimensions. An experimental training pipeline is developed, based on patch-wise learning and advanced data degradation techniques, to simulate real-world video conditions and learn from both local and global structures and details. This helps the model generalize better and maintain stability across varying video content while maintaining the general structure besides the pixel-wise correctness. Two model variants-one larger and one more lightweight-are presented to explore the trade-offs between performance and efficiency. The results demonstrate improved temporal coherence, sharper textures, and fewer visual artifacts compared to traditional single-image methods. This work contributes to the development of practical, learning-based solutions for video enhancement tasks, with potential applications in streaming, gaming, and digital restoration.
format Preprint
id arxiv_https___arxiv_org_abs_2505_10589
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Super-Resolution Generative Adversarial Networks based Video Enhancement
Çetin, Kağan
Akça, Hacer
Gerek, Ömer Nezih
Computer Vision and Pattern Recognition
Artificial Intelligence
Image and Video Processing
I.4.3
This study introduces an enhanced approach to video super-resolution by extending ordinary Single-Image Super-Resolution (SISR) Super-Resolution Generative Adversarial Network (SRGAN) structure to handle spatio-temporal data. While SRGAN has proven effective for single-image enhancement, its design does not account for the temporal continuity required in video processing. To address this, a modified framework that incorporates 3D Non-Local Blocks is proposed, which is enabling the model to capture relationships across both spatial and temporal dimensions. An experimental training pipeline is developed, based on patch-wise learning and advanced data degradation techniques, to simulate real-world video conditions and learn from both local and global structures and details. This helps the model generalize better and maintain stability across varying video content while maintaining the general structure besides the pixel-wise correctness. Two model variants-one larger and one more lightweight-are presented to explore the trade-offs between performance and efficiency. The results demonstrate improved temporal coherence, sharper textures, and fewer visual artifacts compared to traditional single-image methods. This work contributes to the development of practical, learning-based solutions for video enhancement tasks, with potential applications in streaming, gaming, and digital restoration.
title Super-Resolution Generative Adversarial Networks based Video Enhancement
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Image and Video Processing
I.4.3
url https://arxiv.org/abs/2505.10589