The Multi-modal Convolutional Auto-Encoder (MCAE) framework outperformed baseline algorithms in trajectory similarity computation and clustering tasks, achieving a minimum AC value of 0.0465, maximum SI of 0.93003, and minimum CV of 0.1198.
The proposed MCAE framework significantly outperforms baseline algorithms in vessel trajectory similarity computation and clustering tasks.
To enable fast and effective similarity computation between vessel trajectories, this paper proposes a novel Multi-modal Convolutional Auto-Encoder (MCAE) framework for large-scale trajectory representation learning and similarity measurement. The approach converts multi-modal Automatic Identification System (AIS) data—including position(POS), speed over ground (SOG), and course over ground (COG)—into structured image representations. A deep convolutional auto-encoder is employed to automatically learn discriminative features of trajectories. The MCAE model adopts a phased training strategy with a dynamic loss adjustment mechanism to balance learning across different modalities, effectively addressing challenges such as non-uniform sampling, noise interference, and modal heterogeneity in trajectory data. The framework is evaluated using the Adaptive Clustering criterion (AC), silhouette coefficient(SI) and coefficient variation(CV) for comprehensive assessment. Experiments on real-world AIS datasets demonstrate that the proposed method significantly outperforms baseline algorithms—such as Dynamic Time Warping (DTW), Fréchet distance, Convolutional Auto-Encoder (CAE), and t2vec—in trajectory similarity computation and clustering tasks. The method achieves the minimum AC value of 0.0465, the maximum SI of 0.93003 and the minimum CV of 0.1198, indicating stronger robustness and applicability. Furthermore, the research outcomes provide reliable technical support for maritime applications such as vessel monitoring, route planning, and anomaly detection.
Zhang et al. (2026) studied Vessel trajectory similarity measurement (n=5,742). Multi-modal Convolutional Auto-Encoder (MCAE) vs. Dynamic Time Warping (DTW), Fréchet distance, Convolutional Auto-Encoder (CAE), and t2vec was evaluated on Adaptive Clustering criterion (AC), silhouette coefficient (SI), and coefficient variation (CV). The Multi-modal Convolutional Auto-Encoder (MCAE) framework outperformed baseline algorithms in trajectory similarity computation and clustering tasks, achieving a minimum AC value of 0.0465, maximum SI of 0.93003, and minimum CV of 0.1198.