VIRAL: Visual In-Context Reasoning via Analogy in Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910010028589056 |
|---|---|
| author | Li, Zhiwen Duan, Zhongjie Ye, Jinyan Chen, Cen Chen, Daoyuan Li, Yaliang Chen, Yingda |
| author_facet | Li, Zhiwen Duan, Zhongjie Ye, Jinyan Chen, Cen Chen, Daoyuan Li, Yaliang Chen, Yingda |
| contents | Replicating In-Context Learning (ICL) in computer vision remains challenging due to task heterogeneity. We propose \textbf{VIRAL}, a framework that elicits visual reasoning from a pre-trained image editing model by formulating ICL as conditional generation via visual analogy ($x_s : x_t :: x_q : y_q$). We adapt a frozen Diffusion Transformer (DiT) using role-aware multi-image conditioning and introduce a Mixture-of-Experts LoRA to mitigate gradient interference across diverse tasks. Additionally, to bridge the gaps in current visual context datasets, we curate a large-scale dataset spanning perception, restoration, and editing. Experiments demonstrate that VIRAL outperforms existing methods, validating that a unified V-ICL paradigm can handle the majority of visual tasks, including open-domain editing. Our code is available at https://anonymous.4open.science/r/VIRAL-744A |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_03210 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | VIRAL: Visual In-Context Reasoning via Analogy in Diffusion Transformers Li, Zhiwen Duan, Zhongjie Ye, Jinyan Chen, Cen Chen, Daoyuan Li, Yaliang Chen, Yingda Computer Vision and Pattern Recognition Replicating In-Context Learning (ICL) in computer vision remains challenging due to task heterogeneity. We propose \textbf{VIRAL}, a framework that elicits visual reasoning from a pre-trained image editing model by formulating ICL as conditional generation via visual analogy ($x_s : x_t :: x_q : y_q$). We adapt a frozen Diffusion Transformer (DiT) using role-aware multi-image conditioning and introduce a Mixture-of-Experts LoRA to mitigate gradient interference across diverse tasks. Additionally, to bridge the gaps in current visual context datasets, we curate a large-scale dataset spanning perception, restoration, and editing. Experiments demonstrate that VIRAL outperforms existing methods, validating that a unified V-ICL paradigm can handle the majority of visual tasks, including open-domain editing. Our code is available at https://anonymous.4open.science/r/VIRAL-744A |
| title | VIRAL: Visual In-Context Reasoning via Analogy in Diffusion Transformers |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2602.03210 |