AUTV: Creating Underwater Video Datasets with Pixel-wise Annotations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Truong, Quang Trung, Kwan, Wong Yuk, Nguyen, Duc Thanh, Hua, Binh-Son, Yeung, Sai-Kit
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909540654514176
author Truong, Quang Trung
Kwan, Wong Yuk
Nguyen, Duc Thanh
Hua, Binh-Son
Yeung, Sai-Kit
author_facet Truong, Quang Trung
Kwan, Wong Yuk
Nguyen, Duc Thanh
Hua, Binh-Son
Yeung, Sai-Kit
contents Underwater video analysis, hampered by the dynamic marine environment and camera motion, remains a challenging task in computer vision. Existing training-free video generation techniques, learning motion dynamics on the frame-by-frame basis, often produce poor results with noticeable motion interruptions and misaligments. To address these issues, we propose AUTV, a framework for synthesizing marine video data with pixel-wise annotations. We demonstrate the effectiveness of this framework by constructing two video datasets, namely UTV, a real-world dataset comprising 2,000 video-text pairs, and SUTV, a synthetic video dataset including 10,000 videos with segmentation masks for marine objects. UTV provides diverse underwater videos with comprehensive annotations including appearance, texture, camera intrinsics, lighting, and animal behavior. SUTV can be used to improve underwater downstream tasks, which are demonstrated in video inpainting and video object segmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2503_12828
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AUTV: Creating Underwater Video Datasets with Pixel-wise Annotations
Truong, Quang Trung
Kwan, Wong Yuk
Nguyen, Duc Thanh
Hua, Binh-Son
Yeung, Sai-Kit
Computational Engineering, Finance, and Science
Computer Vision and Pattern Recognition
Underwater video analysis, hampered by the dynamic marine environment and camera motion, remains a challenging task in computer vision. Existing training-free video generation techniques, learning motion dynamics on the frame-by-frame basis, often produce poor results with noticeable motion interruptions and misaligments. To address these issues, we propose AUTV, a framework for synthesizing marine video data with pixel-wise annotations. We demonstrate the effectiveness of this framework by constructing two video datasets, namely UTV, a real-world dataset comprising 2,000 video-text pairs, and SUTV, a synthetic video dataset including 10,000 videos with segmentation masks for marine objects. UTV provides diverse underwater videos with comprehensive annotations including appearance, texture, camera intrinsics, lighting, and animal behavior. SUTV can be used to improve underwater downstream tasks, which are demonstrated in video inpainting and video object segmentation.
title AUTV: Creating Underwater Video Datasets with Pixel-wise Annotations
topic Computational Engineering, Finance, and Science
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.12828