UniSER: A Foundation Model for Unified Soft Effects Removal

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jingdong, Zhang, Lingzhi, Liu, Qing, Chiu, Mang Tik, Barnes, Connelly, Wang, Yizhou, You, Haoran, Liu, Xiaoyang, Zhou, Yuqian, Lin, Zhe, Shechtman, Eli, Amirghodsi, Sohrab, Li, Xin, Wang, Wenping, Zhan, Xiaohang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913066470342656
author Zhang, Jingdong
Zhang, Lingzhi
Liu, Qing
Chiu, Mang Tik
Barnes, Connelly
Wang, Yizhou
You, Haoran
Liu, Xiaoyang
Zhou, Yuqian
Lin, Zhe
Shechtman, Eli
Amirghodsi, Sohrab
Li, Xin
Wang, Wenping
Zhan, Xiaohang
author_facet Zhang, Jingdong
Zhang, Lingzhi
Liu, Qing
Chiu, Mang Tik
Barnes, Connelly
Wang, Yizhou
You, Haoran
Liu, Xiaoyang
Zhou, Yuqian
Lin, Zhe
Shechtman, Eli
Amirghodsi, Sohrab
Li, Xin
Wang, Wenping
Zhan, Xiaohang
contents Digital images are often degraded by soft effects such as lens flare, haze, shadows, and reflections, which reduce aesthetics even though the underlying pixels remain partially visible. The prevailing works address these degradations in isolation, developing highly specialized, specialist models that lack scalability and fail to exploit the shared underlying essences of these restoration problems. Meanwhile, although recent large-scale generalist models (e.g., GPT-4o, Flux Kontext, Nano Banana) offer powerful text-driven editing capabilities, they heavily rely on detailed prompts and often fail to achieve robust removal on such fine-grained tasks while preserving the scene's identity. Leveraging the common essence of soft effects, i.e., semi-transparent occlusions, we introduce a foundational versatile model UniSER, capable of addressing diverse degradations caused by soft effects within a single framework. Our methodology centers on curating a massive 3.8M-pair dataset to ensure robustness and generalization, which includes novel, physically-plausible data to fill critical gaps in public benchmarks, and a tailored training pipeline that fine-tunes a Diffusion Transformer to learn robust restoration priors from this diverse data, integrating fine-grained mask and strength controls. This synergistic approach allows UniSER to significantly outperform both specialist and generalist models, achieving robust, high-fidelity restoration in the wild.
format Preprint
id arxiv_https___arxiv_org_abs_2511_14183
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UniSER: A Foundation Model for Unified Soft Effects Removal
Zhang, Jingdong
Zhang, Lingzhi
Liu, Qing
Chiu, Mang Tik
Barnes, Connelly
Wang, Yizhou
You, Haoran
Liu, Xiaoyang
Zhou, Yuqian
Lin, Zhe
Shechtman, Eli
Amirghodsi, Sohrab
Li, Xin
Wang, Wenping
Zhan, Xiaohang
Computer Vision and Pattern Recognition
Digital images are often degraded by soft effects such as lens flare, haze, shadows, and reflections, which reduce aesthetics even though the underlying pixels remain partially visible. The prevailing works address these degradations in isolation, developing highly specialized, specialist models that lack scalability and fail to exploit the shared underlying essences of these restoration problems. Meanwhile, although recent large-scale generalist models (e.g., GPT-4o, Flux Kontext, Nano Banana) offer powerful text-driven editing capabilities, they heavily rely on detailed prompts and often fail to achieve robust removal on such fine-grained tasks while preserving the scene's identity. Leveraging the common essence of soft effects, i.e., semi-transparent occlusions, we introduce a foundational versatile model UniSER, capable of addressing diverse degradations caused by soft effects within a single framework. Our methodology centers on curating a massive 3.8M-pair dataset to ensure robustness and generalization, which includes novel, physically-plausible data to fill critical gaps in public benchmarks, and a tailored training pipeline that fine-tunes a Diffusion Transformer to learn robust restoration priors from this diverse data, integrating fine-grained mask and strength controls. This synergistic approach allows UniSER to significantly outperform both specialist and generalist models, achieving robust, high-fidelity restoration in the wild.
title UniSER: A Foundation Model for Unified Soft Effects Removal
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.14183