Abstract Machine learning models for autism spectrum disorder (ASD) screening based on video analysis often achieve high accuracy when evaluated on the same dataset used for training, but their performance deteriorates sharply when transferred to external datasets. This leads to a limitation of their practical application and reduces the credibility of the results obtained. To address this challenge, we present a simple and reproducible benchmark on two open, privacy-preserving datasets: Multi-Modal ASD dataset (MMASD) and Engagnition. Both datasets were harmonized in a unified metadata table, with the introduction of a common proxy activity label, derived from 2D skeleton dynamics and accelerometry. We compared two simple tabular models, Logistic Regression and XGBoost, under within-dataset independent and identically distributed (IID) and cross-dataset leave-one-dataset-out (LODO) scenarios. We then explicitly quantified the cross-dataset domain shift between MMASD and Engagnition using summary movement features and Wasserstein distances, and visualised the joint feature space with UMAP. Building on this, we evaluated standard domain adaptation baselines (CORrelation ALignment (CORAL) and importance weighting) and a simple PCA-based representation of movement intensity. These analyses show that movement statistics and task structure differ substantially between datasets, basic covariance alignment can partially recover cross-dataset performance for a low-level intensity proxy, whereas naive importance weighting is unstable and collapsing movement descriptors onto a single principal component degrades transfer performance. In addition, minimal TRIPOD-AI reporting elements and PROBAST-AI risk assessment were applied to increase transparency and reproducibility. Keywords: autism spectrum disorder, ASD, video analysis, multimodal dataset, transfer learning, machine learning, ML.
Kurmashev et al. (Wed,) studied this question.