Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pu, Yuandong, Zhuo, Le, Zhu, Kaiwen, Xie, Liangbin, Zhang, Wenlong, Chen, Xiangyu, Gao, Peng, Qiao, Yu, Dong, Chao, Liu, Yihao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916678280937472
author Pu, Yuandong
Zhuo, Le
Zhu, Kaiwen
Xie, Liangbin
Zhang, Wenlong
Chen, Xiangyu
Gao, Peng
Qiao, Yu
Dong, Chao
Liu, Yihao
author_facet Pu, Yuandong
Zhuo, Le
Zhu, Kaiwen
Xie, Liangbin
Zhang, Wenlong
Chen, Xiangyu
Gao, Peng
Qiao, Yu
Dong, Chao
Liu, Yihao
contents We present Lunima-OmniLV (abbreviated as OmniLV), a universal multimodal multi-task framework for low-level vision that addresses over 100 sub-tasks across four major categories: image restoration, image enhancement, weak-semantic dense prediction, and stylization. OmniLV leverages both textual and visual prompts to offer flexible and user-friendly interactions. Built on Diffusion Transformer (DiT)-based generative priors, our framework supports arbitrary resolutions -- achieving optimal performance at 1K resolution -- while preserving fine-grained details and high fidelity. Through extensive experiments, we demonstrate that separately encoding text and visual instructions, combined with co-training using shallow feature control, is essential to mitigate task ambiguity and enhance multi-task generalization. Our findings also reveal that integrating high-level generative tasks into low-level vision models can compromise detail-sensitive restoration. These insights pave the way for more robust and generalizable low-level vision systems.
format Preprint
id arxiv_https___arxiv_org_abs_2504_04903
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
Pu, Yuandong
Zhuo, Le
Zhu, Kaiwen
Xie, Liangbin
Zhang, Wenlong
Chen, Xiangyu
Gao, Peng
Qiao, Yu
Dong, Chao
Liu, Yihao
Computer Vision and Pattern Recognition
Artificial Intelligence
We present Lunima-OmniLV (abbreviated as OmniLV), a universal multimodal multi-task framework for low-level vision that addresses over 100 sub-tasks across four major categories: image restoration, image enhancement, weak-semantic dense prediction, and stylization. OmniLV leverages both textual and visual prompts to offer flexible and user-friendly interactions. Built on Diffusion Transformer (DiT)-based generative priors, our framework supports arbitrary resolutions -- achieving optimal performance at 1K resolution -- while preserving fine-grained details and high fidelity. Through extensive experiments, we demonstrate that separately encoding text and visual instructions, combined with co-training using shallow feature control, is essential to mitigate task ambiguity and enhance multi-task generalization. Our findings also reveal that integrating high-level generative tasks into low-level vision models can compromise detail-sensitive restoration. These insights pave the way for more robust and generalizable low-level vision systems.
title Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2504.04903