MVT: Mask-Grounded Vision-Language Models for Taxonomy-Aligned Land-Cover Tagging

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Siyi, Wang, Kai, Pang, Weicong, Yang, Ruiming, Chen, Ziru, Gao, Renjun, Lau, Alexis Kai Hon, Gu, Dasa, Zhang, Chenchen, Li, Cheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918278087049216
author Chen, Siyi
Wang, Kai
Pang, Weicong
Yang, Ruiming
Chen, Ziru
Gao, Renjun
Lau, Alexis Kai Hon
Gu, Dasa
Zhang, Chenchen
Li, Cheng
author_facet Chen, Siyi
Wang, Kai
Pang, Weicong
Yang, Ruiming
Chen, Ziru
Gao, Renjun
Lau, Alexis Kai Hon
Gu, Dasa
Zhang, Chenchen
Li, Cheng
contents Land-cover understanding in remote sensing increasingly demands class-agnostic systems that generalize across datasets while remaining spatially precise and interpretable. We study a geometry-first discovery-and-interpretation setting under domain shift, where candidate regions are delineated class-agnostically and supervision avoids lexical class names via anonymized identifiers. Complementary to open-set recognition and open-world learning, we focus on coupling class-agnostic mask evidence with taxonomy-grounded scene interpretation, rather than unknown rejection or continual class expansion. We propose MVT, a three-stage framework that (i) extracts boundary-faithful region masks using SAM2 with domain adaptation, (ii) performs mask-grounded semantic tagging and scene description generation via dual-step LoRA fine-tuning of multimodal LLMs, and (iii) evaluates outputs with LLM-as-judge scoring calibrated by stratified expert ratings. On cross-dataset segmentation transfer (train on OpenEarthMap, evaluate on LoveDA), domain-adapted SAM2 improves mask quality; meanwhile, dual-step MLLM fine-tuning yields more accurate taxonomy-aligned tags and more informative mask-grounded scene descriptions.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18693
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MVT: Mask-Grounded Vision-Language Models for Taxonomy-Aligned Land-Cover Tagging
Chen, Siyi
Wang, Kai
Pang, Weicong
Yang, Ruiming
Chen, Ziru
Gao, Renjun
Lau, Alexis Kai Hon
Gu, Dasa
Zhang, Chenchen
Li, Cheng
Computer Vision and Pattern Recognition
Land-cover understanding in remote sensing increasingly demands class-agnostic systems that generalize across datasets while remaining spatially precise and interpretable. We study a geometry-first discovery-and-interpretation setting under domain shift, where candidate regions are delineated class-agnostically and supervision avoids lexical class names via anonymized identifiers. Complementary to open-set recognition and open-world learning, we focus on coupling class-agnostic mask evidence with taxonomy-grounded scene interpretation, rather than unknown rejection or continual class expansion. We propose MVT, a three-stage framework that (i) extracts boundary-faithful region masks using SAM2 with domain adaptation, (ii) performs mask-grounded semantic tagging and scene description generation via dual-step LoRA fine-tuning of multimodal LLMs, and (iii) evaluates outputs with LLM-as-judge scoring calibrated by stratified expert ratings. On cross-dataset segmentation transfer (train on OpenEarthMap, evaluate on LoveDA), domain-adapted SAM2 improves mask quality; meanwhile, dual-step MLLM fine-tuning yields more accurate taxonomy-aligned tags and more informative mask-grounded scene descriptions.
title MVT: Mask-Grounded Vision-Language Models for Taxonomy-Aligned Land-Cover Tagging
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.18693