Any2Any: Unified Arbitrary Modality Translation for Remote Sensing
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866916044262604800 |
|---|---|
| author | Chen, Haoyang Zhang, Jing Wang, Hebaixu Wang, Shiqin Huang, Pohsun Li, Jiayuan Guo, Haonan Wang, Di Wang, Zheng Du, Bo |
| author_facet | Chen, Haoyang Zhang, Jing Wang, Hebaixu Wang, Shiqin Huang, Pohsun Li, Jiayuan Guo, Haonan Wang, Di Wang, Zheng Du, Bo |
| contents | Multi-modal remote sensing imagery provides complementary observations of the same geographic scene, yet such observations are frequently incomplete in practice. Existing cross-modal translation methods treat each modality pair as an independent task, resulting in quadratic complexity and limited generalization to unseen modality combinations. We formulate Any-to-Any translation as inference over a shared latent representation of the scene, where different modalities correspond to partial observations of the same underlying semantics. Based on this formulation, we propose Any2Any, a unified latent diffusion framework that projects heterogeneous inputs into a geometrically aligned latent space. Such structure performs anchored latent regression with a shared backbone, decoupling modality-specific representation learning from semantic mapping. Moreover, lightweight target-specific residual adapters are used to correct systematic latent mismatches without increasing inference complexity. To support learning under sparse but connected supervision, we introduce RST-1M, the first million-scale remote sensing dataset with paired observations across five sensing modalities, providing supervision anchors for any-to-any translation. Experiments across 14 translation tasks show that Any2Any consistently outperforms pairwise translation methods and exhibits strong zero-shot generalization to unseen modality pairs. Code and models are available at https://github.com/MiliLab/Any2Any. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_04114 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Any2Any: Unified Arbitrary Modality Translation for Remote Sensing Chen, Haoyang Zhang, Jing Wang, Hebaixu Wang, Shiqin Huang, Pohsun Li, Jiayuan Guo, Haonan Wang, Di Wang, Zheng Du, Bo Computer Vision and Pattern Recognition Multi-modal remote sensing imagery provides complementary observations of the same geographic scene, yet such observations are frequently incomplete in practice. Existing cross-modal translation methods treat each modality pair as an independent task, resulting in quadratic complexity and limited generalization to unseen modality combinations. We formulate Any-to-Any translation as inference over a shared latent representation of the scene, where different modalities correspond to partial observations of the same underlying semantics. Based on this formulation, we propose Any2Any, a unified latent diffusion framework that projects heterogeneous inputs into a geometrically aligned latent space. Such structure performs anchored latent regression with a shared backbone, decoupling modality-specific representation learning from semantic mapping. Moreover, lightweight target-specific residual adapters are used to correct systematic latent mismatches without increasing inference complexity. To support learning under sparse but connected supervision, we introduce RST-1M, the first million-scale remote sensing dataset with paired observations across five sensing modalities, providing supervision anchors for any-to-any translation. Experiments across 14 translation tasks show that Any2Any consistently outperforms pairwise translation methods and exhibits strong zero-shot generalization to unseen modality pairs. Code and models are available at https://github.com/MiliLab/Any2Any. |
| title | Any2Any: Unified Arbitrary Modality Translation for Remote Sensing |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2603.04114 |