This scientific article provides a comprehensive analysis of the generation of synthetic (artificially created) data for training artificial intelligence (AI) systems and the prospects of this direction. Modern AI models require enormous amounts of data for training, but collecting real data is often expensive, time-consuming, or impossible due to privacy concerns. The research scientifically substantiates the methods of creating artificial data that resembles real data (generative models, simulations), its advantages, and its limitations. The article examines the use of synthetic data in solving privacy problems, eliminating data shortages, and balancing datasets. The scientific novelty of the article lies in demonstrating that synthetic data is becoming an important tool for the development of AI. As a result of the analyses, recommendations are developed regarding the application of these technologies and their reliability
Xudoyorov et al. (Wed,) studied this question.