PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 27, 2026Journal of Marine Science and Engineering0 citationsOpen Access

Research on Small-Sample Data Augmentation and Prediction Method for Ship Equipment Ordering Target Prices Based on GAN and NVP-D Integration

View Full Paper
KLKai LiSSShengxiang SunCZChen Zhu

Key Points

  • This study aims to improve price prediction accuracy for ship equipment orders under small sample constraints.
  • Proposed a method integrating GAN and NVP-D for data augmentation and price prediction.
  • Used Boruta-Lasso for feature selection to reduce model complexity.
  • Conducted experiments with 33 original samples and 24 features using 5-fold cross-validation.
  • Achieved an RMSE of 0.0675, MAE of 0.0510, and R2 of 0.9228 after augmenting to 400 samples.
  • After feature selection, RMSE decreased to 0.0615 and R2 increased to 0.9341.
  • Demonstrated improved prediction accuracy, but validation with larger datasets is needed.

Abstract

To address the problem in predicting target prices for ship equipment orders where small sample sizes, high feature dimensions, and strong business constraints lead traditional models to overfit and have insufficient generalization ability, a method combining GAN and NVP-D for small-sample data augmentation and price prediction is proposed. This method integrates the advantages of adversarial training in Generative Adversarial Networks (GAN) with the explicit density estimation and stable training characteristics of Normalizing Flow NVP-D. By using dual-weight collaborative optimization of the objective function, it alleviates gradient vanishing and mode collapse, generating high-quality virtual samples that closely follow the real data distribution. Redundant features are removed using Boruta-Lasso joint feature selection to reduce model complexity. CatBoost is employed as the prediction model to complete price estimation. Experiments were conducted on a ship equipment dataset with 33 original samples and 24 features, strictly following the standard procedure of augmentation only within the training set and 5-fold cross-validation. Compared with NVP, NVP-G, MAF, traditional GAN, and Mixup methods, the results show that the proposed integrated model achieves optimal performance when augmenting 400 samples, with an RMSE of 0.0675, MAE of 0.0510, and R2 of 0.9228. After feature selection, prediction accuracy further improves, with RMSE decreasing to 0.0615 and R2 increasing to 0.9341. Limited by the scale of the original samples, the statistical robustness and cross-dataset generalization capability of this method still need validation with larger datasets. However, under the current small-sample constraints, it can effectively alleviate modeling bottlenecks and provide high-precision support for equipment procurement argumentation, budget preparation, and cost control in stages.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2026) studied this question.

synapsesocial.com/papers/6a168b430c924ddd1bd5a2aahttps://doi.org/10.3390/jmse14100923
Ask AI
Helpful
Bookmark
Share
View Full Paper