Key points are not available for this paper at this time.
Utilizing spatiotemporal features in massive amounts of trajectory data to identify the operation mode of agricultural machinery trajectories is a key task in precision agriculture. Most of the previous studies focuses narrowly on single-perspective feature extraction, neglecting comprehensive spatiotemporal information in trajectory data. To improve the accuracy of the task, this paper proposes a model called STGMAE. First, we propose a multilevel feature extraction method (MFE), which extracts motion and statistical features from the initial features via a motion feature extractor and a sliding time window, and then uses the spectral feature module (SFM) to capture the spectral information, which improves the representation of trajectory data. Next, to prevent information loss in long-range encoding, we design a pre-training network with serial encoding and parallel decoding. Specifically, the data are first modeled globally interactively via a multiscale wavelet projector (WMP), and then enter an adaptive graph isomorphic neural network (AGIN). In AGIN, semi-adaptive masked Laplace operator (SAMLO) is used to capture the correlation information between trajectory points, and then a passing mechanism is used to address the homogeneous relationships between trajectory points and heterogeneous relationships between trajectory graphs. Then, the original feature and graph structure are reconstructed from the two encoding nodes to realize self-supervised training. Eventually, we use the pre-trained weights on the real trajectory samples provided by the Key Laboratory of Agricultural Machinery Monitoring and Big Data Application, Ministry of Agriculture and Rural Affairs, which contains 219 trajectory samples (2,250,693 trajectory points). The experimental results show that for the paddy, corn, and wheat harvesting trajectory datasets, our model accuracies are 95.50 %, 95.32 %, and 95.36 %, respectively, and the F1 scores are 94.54 %, 92.09 %, and 93.79 %, respectively. Compared with existing state-of-the-art methods, our method achieves accuracies of 5.75 %, 4.47 %, and 5.03 % and F1 scores of 7.26 %, 4.85 %, and 6.65 %, respectively. • A model using pre-training spatiotemporal graph masked autoencoder was proposed. • Multilevel feature extraction method was used for feature enhancement. • A structure based on serial encoding and parallel decoding was proposed. • Fine-tuning on large-scale data using pre-trained weights.
Chen et al. (Sat,) studied this question.