Accurate short-term metro passenger flow forecasting plays a key role in urban transit management, supporting train scheduling, crowd control, and operational planning. Jointly modeling station-level inflow/outflow (IO) and inter-station origin–destination flows (OD/DO) has proven effective for improving prediction accuracy, as it allows the model to leverage dependencies across different flow granularities. However, effectively exploiting such dependencies remains nontrivial. Station-level intensity (IO) and inter-station migration patterns (OD/DO) differ substantially in both representation and dynamics, and the dependencies between them are inherently directional and uneven. As a result, commonly used parameter-sharing mechanisms in multi-task learning are often insufficient to capture informative cross-task interactions. To address this issue, we propose CATI (Cross-Attention-based Task Interaction), a unified framework for joint multi-granular metro flow forecasting. CATI first learns task-specific spatiotemporal representations for IO, OD, and DO flows, and then introduces directed cross-attention with Gated Residual Fusion to model selective and asymmetric interactions across tasks. In addition, an aggregation-consistency regularization is employed to maintain structural coherence between station-level and inter-station predictions. Experiments on real-world metro datasets from Hangzhou and Shanghai show that CATI consistently outperforms strong baselines across multiple prediction horizons and tasks. Further analysis indicates that the model learns adaptive attention patterns, task-dependent gating behaviors, and controlled interaction strengths, which together explain its improved performance. These results suggest that explicitly modeling asymmetric cross-task interactions is important for multi-granular spatiotemporal forecasting in metro systems.
Yang et al. (Fri,) studied this question.