Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866916678280937472 |
|---|---|
| author | Pu, Yuandong Zhuo, Le Zhu, Kaiwen Xie, Liangbin Zhang, Wenlong Chen, Xiangyu Gao, Peng Qiao, Yu Dong, Chao Liu, Yihao |
| author_facet | Pu, Yuandong Zhuo, Le Zhu, Kaiwen Xie, Liangbin Zhang, Wenlong Chen, Xiangyu Gao, Peng Qiao, Yu Dong, Chao Liu, Yihao |
| contents | We present Lunima-OmniLV (abbreviated as OmniLV), a universal multimodal multi-task framework for low-level vision that addresses over 100 sub-tasks across four major categories: image restoration, image enhancement, weak-semantic dense prediction, and stylization. OmniLV leverages both textual and visual prompts to offer flexible and user-friendly interactions. Built on Diffusion Transformer (DiT)-based generative priors, our framework supports arbitrary resolutions -- achieving optimal performance at 1K resolution -- while preserving fine-grained details and high fidelity. Through extensive experiments, we demonstrate that separately encoding text and visual instructions, combined with co-training using shallow feature control, is essential to mitigate task ambiguity and enhance multi-task generalization. Our findings also reveal that integrating high-level generative tasks into low-level vision models can compromise detail-sensitive restoration. These insights pave the way for more robust and generalizable low-level vision systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_04903 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision Pu, Yuandong Zhuo, Le Zhu, Kaiwen Xie, Liangbin Zhang, Wenlong Chen, Xiangyu Gao, Peng Qiao, Yu Dong, Chao Liu, Yihao Computer Vision and Pattern Recognition Artificial Intelligence We present Lunima-OmniLV (abbreviated as OmniLV), a universal multimodal multi-task framework for low-level vision that addresses over 100 sub-tasks across four major categories: image restoration, image enhancement, weak-semantic dense prediction, and stylization. OmniLV leverages both textual and visual prompts to offer flexible and user-friendly interactions. Built on Diffusion Transformer (DiT)-based generative priors, our framework supports arbitrary resolutions -- achieving optimal performance at 1K resolution -- while preserving fine-grained details and high fidelity. Through extensive experiments, we demonstrate that separately encoding text and visual instructions, combined with co-training using shallow feature control, is essential to mitigate task ambiguity and enhance multi-task generalization. Our findings also reveal that integrating high-level generative tasks into low-level vision models can compromise detail-sensitive restoration. These insights pave the way for more robust and generalizable low-level vision systems. |
| title | Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2504.04903 |