Goal-Conditioned Reinforcement Learning for Data-Driven Maritime Navigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vaidheeswaran, Vaishnav, Jayakody, Dilith, Mulay, Samruddhi, Lo, Anand, Alam, Md Mahbub, Spadon, Gabriel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911133581967360
author Vaidheeswaran, Vaishnav
Jayakody, Dilith
Mulay, Samruddhi
Lo, Anand
Alam, Md Mahbub
Spadon, Gabriel
author_facet Vaidheeswaran, Vaishnav
Jayakody, Dilith
Mulay, Samruddhi
Lo, Anand
Alam, Md Mahbub
Spadon, Gabriel
contents Routing vessels through narrow and dynamic waterways is challenging due to changing environmental conditions and operational constraints. Existing vessel-routing studies typically fail to generalize across multiple origin-destination pairs and do not exploit large-scale, data-driven traffic graphs. In this paper, we propose a reinforcement learning solution for big maritime data that can learn to find a route across multiple origin-destination pairs while adapting to different hexagonal grid resolutions. Agents learn to select direction and speed under continuous observations in a multi-discrete action space. A reward function balances fuel efficiency, travel time, wind resistance, and route diversity, using an Automatic Identification System (AIS)-derived traffic graph with ERA5 wind fields. The approach is demonstrated in the Gulf of St. Lawrence, one of the largest estuaries in the world. We evaluate configurations that combine Proximal Policy Optimization with recurrent networks, invalid-action masking, and exploration strategies. Our experiments demonstrate that action masking yields a clear improvement in policy performance and that supplementing penalty-only feedback with positive shaping rewards produces additional gains.
format Preprint
id arxiv_https___arxiv_org_abs_2509_01838
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Goal-Conditioned Reinforcement Learning for Data-Driven Maritime Navigation
Vaidheeswaran, Vaishnav
Jayakody, Dilith
Mulay, Samruddhi
Lo, Anand
Alam, Md Mahbub
Spadon, Gabriel
Machine Learning
Artificial Intelligence
Routing vessels through narrow and dynamic waterways is challenging due to changing environmental conditions and operational constraints. Existing vessel-routing studies typically fail to generalize across multiple origin-destination pairs and do not exploit large-scale, data-driven traffic graphs. In this paper, we propose a reinforcement learning solution for big maritime data that can learn to find a route across multiple origin-destination pairs while adapting to different hexagonal grid resolutions. Agents learn to select direction and speed under continuous observations in a multi-discrete action space. A reward function balances fuel efficiency, travel time, wind resistance, and route diversity, using an Automatic Identification System (AIS)-derived traffic graph with ERA5 wind fields. The approach is demonstrated in the Gulf of St. Lawrence, one of the largest estuaries in the world. We evaluate configurations that combine Proximal Policy Optimization with recurrent networks, invalid-action masking, and exploration strategies. Our experiments demonstrate that action masking yields a clear improvement in policy performance and that supplementing penalty-only feedback with positive shaping rewards produces additional gains.
title Goal-Conditioned Reinforcement Learning for Data-Driven Maritime Navigation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.01838