Randomized trial demonstrates effective subsampling strategy for multivariate longitudinal data, suggesting better computational efficiency.
Longitudinal studies routinely generate large-scale multivariate data in which multiple responses are measured repeatedly for each subject over time. Fitting the popular class of linear mixed-effects models to such data and obtaining maximum likelihood estimates for the model parameters can be computationally prohibitive when the number of subjects m is large. This is because the likelihood function must be evaluated over the full dataset at every iteration of the parameter estimation algorithm, so that the cost of each iteration grows with both the number of subjects m and the number of repeated measurements ni per subject; for very large m, repeatedly evaluating the likelihood over all subjects quickly becomes infeasible. To alleviate this problem, we propose an optimal subsampling framework for multivariate longitudinal data within the linear mixed-effects model. This framework reduces the computational burden by fitting the model on a small carefully selected subset of subjects rather than the full sample, so that the per-iteration cost depends on the much smaller subsample size instead of the total number of subjects m. We first establish the conditional asymptotic distribution of the subsample estimator under general subsampling probabilities and then show that the optimal subsampling probabilities—those minimizing the asymptotic mean squared error (AMSE) of the subsample estimator—admit a closed-form expression analogous to the A-optimality criterion from the theory of optimal experimental design. We further derive L-optimal subsampling probabilities for settings in which only a linear transformation of the full parameter vector is of primary interest. Because the optimal probabilities depend on the unknown full-data maximum likelihood estimator, we develop a practical two-step algorithm in which a small pilot subsample is first drawn uniformly to obtain initial estimates, which are then used to compute approximated optimal probabilities for a second larger subsample. Simulation experiments demonstrate that the proposed optimal subsampling method improves upon commonly used uniform subsampling. We provide practical recommendations to guide practitioners in choosing between uniform and optimal subsampling.
No takes yet. Share an insight, caveat, or question.
Wang et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: