Edge-Efficient Image Restoration: Transformer Distillation into State-Space Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Miriyala, Srinivas Soumitri, Vajrala, Sowmya, Kodavanti, Sravanth, Rajendiran, Vikram Nelvoy, Allur, Sharan Kumar
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917459181699072
author Miriyala, Srinivas Soumitri
Vajrala, Sowmya
Kodavanti, Sravanth
Rajendiran, Vikram Nelvoy
Allur, Sharan Kumar
author_facet Miriyala, Srinivas Soumitri
Vajrala, Sowmya
Kodavanti, Sravanth
Rajendiran, Vikram Nelvoy
Allur, Sharan Kumar
contents We propose a modular framework for hybrid image restoration that integrates transformer and state-space model (SSM) blocks with a focus on improving runtime efficiency on edge hardware. While transformers provide strong global modeling through self-attention, their attention kernels incur substantial latency on mobile devices, especially for high-resolution inputs. In contrast, SSMs such as Mamba offer lineartime sequence modeling with lower runtime overhead but may underperform on fine grained restoration tasks. To balance accuracy and efficiency, we train lightweight SSM blocks as feature-distilled surrogates of transformer blocks and use them to construct hybrid U-Net-style architectures. To automatically discover effective block combinations, we introduce Efficient Network Search (ENS), a multi-objective search strategy that selects task-specific hybrid configurations from pre-aligned components. ENS optimizes restoration quality while penalizing transformer usage, serving as a lightweight proxy for latency and enabling architecture discovery without repeated hardware profiling. On a Snapdragon 8 Elite CPU, the Restormer baseline requires 10119.52 ms for inference. In contrast, ENS-discovered hybrids significantly reduce runtime: ENS-Deblurring runs in 2973 ms (3.4x faster), ENS-Deraining in 5816 ms (1.74x faster), and ENS-Denoising in 8666 ms (1.17x faster), while maintaining competitive restoration quality.
format Preprint
id arxiv_https___arxiv_org_abs_2605_02794
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Edge-Efficient Image Restoration: Transformer Distillation into State-Space Models
Miriyala, Srinivas Soumitri
Vajrala, Sowmya
Kodavanti, Sravanth
Rajendiran, Vikram Nelvoy
Allur, Sharan Kumar
Computer Vision and Pattern Recognition
We propose a modular framework for hybrid image restoration that integrates transformer and state-space model (SSM) blocks with a focus on improving runtime efficiency on edge hardware. While transformers provide strong global modeling through self-attention, their attention kernels incur substantial latency on mobile devices, especially for high-resolution inputs. In contrast, SSMs such as Mamba offer lineartime sequence modeling with lower runtime overhead but may underperform on fine grained restoration tasks. To balance accuracy and efficiency, we train lightweight SSM blocks as feature-distilled surrogates of transformer blocks and use them to construct hybrid U-Net-style architectures. To automatically discover effective block combinations, we introduce Efficient Network Search (ENS), a multi-objective search strategy that selects task-specific hybrid configurations from pre-aligned components. ENS optimizes restoration quality while penalizing transformer usage, serving as a lightweight proxy for latency and enabling architecture discovery without repeated hardware profiling. On a Snapdragon 8 Elite CPU, the Restormer baseline requires 10119.52 ms for inference. In contrast, ENS-discovered hybrids significantly reduce runtime: ENS-Deblurring runs in 2973 ms (3.4x faster), ENS-Deraining in 5816 ms (1.74x faster), and ENS-Denoising in 8666 ms (1.17x faster), while maintaining competitive restoration quality.
title Edge-Efficient Image Restoration: Transformer Distillation into State-Space Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.02794