Measuring the State of Open Science in Transportation Using Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ji, Junyi, Lu, Ruth, Belkessa, Linda, Wang, Liming, Varotto, Silvia, Dong, Yongqi, Saunier, Nicolas, Ameli, Mostafa, Macfarlane, Gregory S., Madadi, Bahman, Wu, Cathy
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912836748312576
author Ji, Junyi
Lu, Ruth
Belkessa, Linda
Wang, Liming
Varotto, Silvia
Dong, Yongqi
Saunier, Nicolas
Ameli, Mostafa
Macfarlane, Gregory S.
Madadi, Bahman
Wu, Cathy
author_facet Ji, Junyi
Lu, Ruth
Belkessa, Linda
Wang, Liming
Varotto, Silvia
Dong, Yongqi
Saunier, Nicolas
Ameli, Mostafa
Macfarlane, Gregory S.
Madadi, Bahman
Wu, Cathy
contents Open science initiatives have strengthened scientific integrity and accelerated research progress across many fields, but the state of their practice within transportation research remains under-investigated. Key features of open science, defined here as data and code availability, are difficult to extract due to the inherent complexity of the field. Previous work has either been limited to small-scale studies due to the labor-intensive nature of manual analysis or has relied on large-scale bibliometric approaches that sacrifice contextual richness. This paper introduces an automatic and scalable feature-extraction pipeline to measure data and code availability in transportation research. We employ Large Language Models (LLMs) for this task and validate their performance against a manually curated dataset and through an inter-rater agreement analysis. We applied this pipeline to examine 10,724 research articles published in the Transportation Research Part series of journals between 2019 and 2024. Our analysis found that only 5% of quantitative papers shared a code repository, 4% of quantitative papers shared a data repository, and about 3% of papers shared both, with trends differing across journals, topics, and geographic regions. We found no significant difference in citation counts or review duration between papers that provided data and code and those that did not, suggesting a misalignment between open science efforts and traditional academic metrics. Consequently, encouraging these practices will likely require structural interventions from journals and funding agencies to supplement the lack of direct author incentives. The pipeline developed in this study can be readily scaled to other journals, representing a critical step toward the automated measurement and monitoring of open science practices in transportation research.
format Preprint
id arxiv_https___arxiv_org_abs_2601_14429
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Measuring the State of Open Science in Transportation Using Large Language Models
Ji, Junyi
Lu, Ruth
Belkessa, Linda
Wang, Liming
Varotto, Silvia
Dong, Yongqi
Saunier, Nicolas
Ameli, Mostafa
Macfarlane, Gregory S.
Madadi, Bahman
Wu, Cathy
Digital Libraries
Artificial Intelligence
Computers and Society
Emerging Technologies
Open science initiatives have strengthened scientific integrity and accelerated research progress across many fields, but the state of their practice within transportation research remains under-investigated. Key features of open science, defined here as data and code availability, are difficult to extract due to the inherent complexity of the field. Previous work has either been limited to small-scale studies due to the labor-intensive nature of manual analysis or has relied on large-scale bibliometric approaches that sacrifice contextual richness. This paper introduces an automatic and scalable feature-extraction pipeline to measure data and code availability in transportation research. We employ Large Language Models (LLMs) for this task and validate their performance against a manually curated dataset and through an inter-rater agreement analysis. We applied this pipeline to examine 10,724 research articles published in the Transportation Research Part series of journals between 2019 and 2024. Our analysis found that only 5% of quantitative papers shared a code repository, 4% of quantitative papers shared a data repository, and about 3% of papers shared both, with trends differing across journals, topics, and geographic regions. We found no significant difference in citation counts or review duration between papers that provided data and code and those that did not, suggesting a misalignment between open science efforts and traditional academic metrics. Consequently, encouraging these practices will likely require structural interventions from journals and funding agencies to supplement the lack of direct author incentives. The pipeline developed in this study can be readily scaled to other journals, representing a critical step toward the automated measurement and monitoring of open science practices in transportation research.
title Measuring the State of Open Science in Transportation Using Large Language Models
topic Digital Libraries
Artificial Intelligence
Computers and Society
Emerging Technologies
url https://arxiv.org/abs/2601.14429