CyPortQA: Benchmarking Multimodal Large Language Models for Cyclone Preparedness in Port Operation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kuai, Chenchen, Wu, Chenhao, Zhou, Yang, Wang, Xiubin Bruce, Yang, Tianbao, Tu, Zhengzhong, Li, Zihao, Zhang, Yunlong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914157589168128
author Kuai, Chenchen
Wu, Chenhao
Zhou, Yang
Wang, Xiubin Bruce
Yang, Tianbao
Tu, Zhengzhong
Li, Zihao
Zhang, Yunlong
author_facet Kuai, Chenchen
Wu, Chenhao
Zhou, Yang
Wang, Xiubin Bruce
Yang, Tianbao
Tu, Zhengzhong
Li, Zihao
Zhang, Yunlong
contents As tropical cyclones intensify and track forecasts become increasingly uncertain, U.S. ports face heightened supply-chain risk under extreme weather conditions. Port operators need to rapidly synthesize diverse multimodal forecast products, such as probabilistic wind maps, track cones, and official advisories, into clear, actionable guidance as cyclones approach. Multimodal large language models (MLLMs) offer a powerful means to integrate these heterogeneous data sources alongside broader contextual knowledge, yet their accuracy and reliability in the specific context of port cyclone preparedness have not been rigorously evaluated. To fill this gap, we introduce CyPortQA, the first multimodal benchmark tailored to port operations under cyclone threat. CyPortQA assembles 2,917 realworld disruption scenarios from 2015 through 2023, spanning 145 U.S. principal ports and 90 named storms. Each scenario fuses multisource data (i.e., tropical cyclone products, port operational impact records, and port condition bulletins) and is expanded through an automated pipeline into 117,178 structured question answer pairs. Using this benchmark, we conduct extensive experiments on diverse MLLMs, including both open-source and proprietary model. MLLMs demonstrate great potential in situation understanding but still face considerable challenges in reasoning tasks, including potential impact estimation and decision reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15846
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CyPortQA: Benchmarking Multimodal Large Language Models for Cyclone Preparedness in Port Operation
Kuai, Chenchen
Wu, Chenhao
Zhou, Yang
Wang, Xiubin Bruce
Yang, Tianbao
Tu, Zhengzhong
Li, Zihao
Zhang, Yunlong
Computation and Language
As tropical cyclones intensify and track forecasts become increasingly uncertain, U.S. ports face heightened supply-chain risk under extreme weather conditions. Port operators need to rapidly synthesize diverse multimodal forecast products, such as probabilistic wind maps, track cones, and official advisories, into clear, actionable guidance as cyclones approach. Multimodal large language models (MLLMs) offer a powerful means to integrate these heterogeneous data sources alongside broader contextual knowledge, yet their accuracy and reliability in the specific context of port cyclone preparedness have not been rigorously evaluated. To fill this gap, we introduce CyPortQA, the first multimodal benchmark tailored to port operations under cyclone threat. CyPortQA assembles 2,917 realworld disruption scenarios from 2015 through 2023, spanning 145 U.S. principal ports and 90 named storms. Each scenario fuses multisource data (i.e., tropical cyclone products, port operational impact records, and port condition bulletins) and is expanded through an automated pipeline into 117,178 structured question answer pairs. Using this benchmark, we conduct extensive experiments on diverse MLLMs, including both open-source and proprietary model. MLLMs demonstrate great potential in situation understanding but still face considerable challenges in reasoning tasks, including potential impact estimation and decision reasoning.
title CyPortQA: Benchmarking Multimodal Large Language Models for Cyclone Preparedness in Port Operation
topic Computation and Language
url https://arxiv.org/abs/2508.15846