MVT: Mask-Grounded Vision-Language Models for Taxonomy-Aligned Land-Cover Tagging
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918278087049216 |
|---|---|
| author | Chen, Siyi Wang, Kai Pang, Weicong Yang, Ruiming Chen, Ziru Gao, Renjun Lau, Alexis Kai Hon Gu, Dasa Zhang, Chenchen Li, Cheng |
| author_facet | Chen, Siyi Wang, Kai Pang, Weicong Yang, Ruiming Chen, Ziru Gao, Renjun Lau, Alexis Kai Hon Gu, Dasa Zhang, Chenchen Li, Cheng |
| contents | Land-cover understanding in remote sensing increasingly demands class-agnostic systems that generalize across datasets while remaining spatially precise and interpretable. We study a geometry-first discovery-and-interpretation setting under domain shift, where candidate regions are delineated class-agnostically and supervision avoids lexical class names via anonymized identifiers. Complementary to open-set recognition and open-world learning, we focus on coupling class-agnostic mask evidence with taxonomy-grounded scene interpretation, rather than unknown rejection or continual class expansion. We propose MVT, a three-stage framework that (i) extracts boundary-faithful region masks using SAM2 with domain adaptation, (ii) performs mask-grounded semantic tagging and scene description generation via dual-step LoRA fine-tuning of multimodal LLMs, and (iii) evaluates outputs with LLM-as-judge scoring calibrated by stratified expert ratings. On cross-dataset segmentation transfer (train on OpenEarthMap, evaluate on LoveDA), domain-adapted SAM2 improves mask quality; meanwhile, dual-step MLLM fine-tuning yields more accurate taxonomy-aligned tags and more informative mask-grounded scene descriptions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_18693 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MVT: Mask-Grounded Vision-Language Models for Taxonomy-Aligned Land-Cover Tagging Chen, Siyi Wang, Kai Pang, Weicong Yang, Ruiming Chen, Ziru Gao, Renjun Lau, Alexis Kai Hon Gu, Dasa Zhang, Chenchen Li, Cheng Computer Vision and Pattern Recognition Land-cover understanding in remote sensing increasingly demands class-agnostic systems that generalize across datasets while remaining spatially precise and interpretable. We study a geometry-first discovery-and-interpretation setting under domain shift, where candidate regions are delineated class-agnostically and supervision avoids lexical class names via anonymized identifiers. Complementary to open-set recognition and open-world learning, we focus on coupling class-agnostic mask evidence with taxonomy-grounded scene interpretation, rather than unknown rejection or continual class expansion. We propose MVT, a three-stage framework that (i) extracts boundary-faithful region masks using SAM2 with domain adaptation, (ii) performs mask-grounded semantic tagging and scene description generation via dual-step LoRA fine-tuning of multimodal LLMs, and (iii) evaluates outputs with LLM-as-judge scoring calibrated by stratified expert ratings. On cross-dataset segmentation transfer (train on OpenEarthMap, evaluate on LoveDA), domain-adapted SAM2 improves mask quality; meanwhile, dual-step MLLM fine-tuning yields more accurate taxonomy-aligned tags and more informative mask-grounded scene descriptions. |
| title | MVT: Mask-Grounded Vision-Language Models for Taxonomy-Aligned Land-Cover Tagging |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2509.18693 |