Key points are not available for this paper at this time.
Abstract Recent advances in artificial intelligence offer considerable promise for enhancing industrial computational simulation workflows, yet practical adoption is limited by a shortage of high-quality, real-world datasets. Most existing CFD-related datasets focus on a single phase of the simulation lifecycle, typically post-processed flow snapshots, and lack the upstream information needed to train models that can exploit the full pipeline. In this work, we present three such datasets, collected from the China Ship Scientific Research Center during operational industrial workflows. The data cover three marine vessels-two tanker ships and a submarine-whose low-speed, high-block-coefficient hull forms produce complex, separated stern flows. Uniquely, and in contrast to conventional repositories, our dataset spans the complete CFD lifecycle: pre-processing (experimental documentation and geometries), solving (the linear systems A x = b resulting from discretisation), and post-processing (numerical solutions and computational meshes). By assembling these interdependent stages into a single, coherent resource, the dataset is designed to support the development of robust AI models that can learn across phases, while also facilitating research in solver performance evaluation, data management, and flow visualization. Beyond AI, the dataset serves a wider range of research topics, such as benchmark-driven solver evaluation, data-management studies, and flow-field visualization, enhancing its overall utility and impact.
Wang et al. (Sat,) studied this question.