Multimodal Trajectory Representation Learning for Travel Time Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Zhi, Hu, Xuyuan, Han, Xiao, Dai, Zhehao, Deng, Zhaolin, Shen, Guojiang, Kong, Xiangjie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915754532667392
author Liu, Zhi
Hu, Xuyuan
Han, Xiao
Dai, Zhehao
Deng, Zhaolin
Shen, Guojiang
Kong, Xiangjie
author_facet Liu, Zhi
Hu, Xuyuan
Han, Xiao
Dai, Zhehao
Deng, Zhaolin
Shen, Guojiang
Kong, Xiangjie
contents Accurate travel time estimation (TTE) plays a crucial role in intelligent transportation systems. However, it remains challenging due to heterogeneous data sources and complex traffic dynamics. Moreover, traditional approaches typically convert trajectory data into fixed-length representations. This overlooks the inherent variability of real-world motion patterns, often resulting in information loss and redundancy. To address these challenges, this paper introduces the Multimodal Dynamic Trajectory Integration (MDTI) framework--a novel multimodal trajectory representation learning approach that integrates GPS sequences, grid trajectories, and road network constraints to enhance the performance of TTE. MDTI employs modality-specific encoders and a multimodal fusion module to capture complementary spatial, temporal, and topological semantics, while a dynamic trajectory modeling mechanism adaptively regulates information density for trajectories of varying lengths. Two self-supervised pretraining objectives, named contrastive alignment and masked language modeling, further strengthen multimodal consistency and contextual understanding. Extensive experiments on three real-world datasets demonstrate that MDTI consistently outperforms state-of-the-art baselines, confirming its robustness and strong generalization abilities. The code is publicly available at: https://github.com/City-Computing/MDTI.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05840
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal Trajectory Representation Learning for Travel Time Estimation
Liu, Zhi
Hu, Xuyuan
Han, Xiao
Dai, Zhehao
Deng, Zhaolin
Shen, Guojiang
Kong, Xiangjie
Machine Learning
Accurate travel time estimation (TTE) plays a crucial role in intelligent transportation systems. However, it remains challenging due to heterogeneous data sources and complex traffic dynamics. Moreover, traditional approaches typically convert trajectory data into fixed-length representations. This overlooks the inherent variability of real-world motion patterns, often resulting in information loss and redundancy. To address these challenges, this paper introduces the Multimodal Dynamic Trajectory Integration (MDTI) framework--a novel multimodal trajectory representation learning approach that integrates GPS sequences, grid trajectories, and road network constraints to enhance the performance of TTE. MDTI employs modality-specific encoders and a multimodal fusion module to capture complementary spatial, temporal, and topological semantics, while a dynamic trajectory modeling mechanism adaptively regulates information density for trajectories of varying lengths. Two self-supervised pretraining objectives, named contrastive alignment and masked language modeling, further strengthen multimodal consistency and contextual understanding. Extensive experiments on three real-world datasets demonstrate that MDTI consistently outperforms state-of-the-art baselines, confirming its robustness and strong generalization abilities. The code is publicly available at: https://github.com/City-Computing/MDTI.
title Multimodal Trajectory Representation Learning for Travel Time Estimation
topic Machine Learning
url https://arxiv.org/abs/2510.05840