VDD: Varied Drone Dataset for Semantic Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cai, Wenxiao, Jin, Ke, Hou, Jinyan, Guo, Cong, Wu, Letian, Yang, Wankou
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909556404125696
author Cai, Wenxiao
Jin, Ke
Hou, Jinyan
Guo, Cong
Wu, Letian
Yang, Wankou
author_facet Cai, Wenxiao
Jin, Ke
Hou, Jinyan
Guo, Cong
Wu, Letian
Yang, Wankou
contents Semantic segmentation of drone images is critical for various aerial vision tasks as it provides essential semantic details to understand scenes on the ground. Ensuring high accuracy of semantic segmentation models for drones requires access to diverse, large-scale, and high-resolution datasets, which are often scarce in the field of aerial image processing. While existing datasets typically focus on urban scenes and are relatively small, our Varied Drone Dataset (VDD) addresses these limitations by offering a large-scale, densely labeled collection of 400 high-resolution images spanning 7 classes. This dataset features various scenes in urban, industrial, rural, and natural areas, captured from different camera angles and under diverse lighting conditions. We also make new annotations to UDD and UAVid, integrating them under VDD annotation standards, to create the Integrated Drone Dataset (IDD). We train seven state-of-the-art models on drone datasets as baselines. It's expected that our dataset will generate considerable interest in drone image segmentation and serve as a foundation for other drone vision tasks. Datasets are publicly available at \href{our website}{https://github.com/RussRobin/VDD}.
format Preprint
id arxiv_https___arxiv_org_abs_2305_13608
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle VDD: Varied Drone Dataset for Semantic Segmentation
Cai, Wenxiao
Jin, Ke
Hou, Jinyan
Guo, Cong
Wu, Letian
Yang, Wankou
Computer Vision and Pattern Recognition
Semantic segmentation of drone images is critical for various aerial vision tasks as it provides essential semantic details to understand scenes on the ground. Ensuring high accuracy of semantic segmentation models for drones requires access to diverse, large-scale, and high-resolution datasets, which are often scarce in the field of aerial image processing. While existing datasets typically focus on urban scenes and are relatively small, our Varied Drone Dataset (VDD) addresses these limitations by offering a large-scale, densely labeled collection of 400 high-resolution images spanning 7 classes. This dataset features various scenes in urban, industrial, rural, and natural areas, captured from different camera angles and under diverse lighting conditions. We also make new annotations to UDD and UAVid, integrating them under VDD annotation standards, to create the Integrated Drone Dataset (IDD). We train seven state-of-the-art models on drone datasets as baselines. It's expected that our dataset will generate considerable interest in drone image segmentation and serve as a foundation for other drone vision tasks. Datasets are publicly available at \href{our website}{https://github.com/RussRobin/VDD}.
title VDD: Varied Drone Dataset for Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2305.13608