This study introduces an integrated empirical–computational framework for predicting antioxidant activity using heterogeneous biochemical data. Experimental measurements of total phenolic content (TPC), total flavonoid content (TFC), and half-maximal inhibitory concentration (IC 50 ) from methanol, ethanol, and acetone extracts of Trigonella foenum-graecum L. and Momordica charantia L. were combined with in silico parametric simulations spanning fourteen solvent environments to construct a structured feature space capturing solvent–bioactivity interactions. Polynomial and ensemble models successfully learned nonlinear relationships (coefficient of determination, R 2 = 0 . 73 in Phase I), and Monte Carlo perturbation analysis demonstrated predictive stability ( ± 12 % ). In Phase II, 54 empirical measurements were integrated with 6 literature-derived records ( n = 60 ) under group-stratified train/test partitioning and log-transformation of IC 50 to ensure honest evaluation across heterogeneous data sources. Random Forest (RF) achieved the robust predictive performance ( R 2 = 0 . 931 , Adj. R 2 = 0 . 896 , root mean square error (RMSE) = 75.44 μ g mL −1 ), with model validity confirmed by Y-randomisation (mean permuted R 2 = − 0 . 146 ) and bootstrap resampling (95% confidence interval (CI): 0 . 887 , 0 . 977 ). Feature importance and SHapley Additive exPlanations (SHAP) analysis identified TPC and TFC as the dominant predictors of antioxidant response, with solvent polarity contributing secondary explanatory influence, consistent with observed correlations ( r TPC = − 0 . 70 , r TFC = − 0 . 31 ). This work provides a reproducible, data-driven computational pipeline for modeling antioxidant behavior across diverse experimental conditions and demonstrates a statistically validated prediction strategy applicable to phytochemical and biochemical systems.
Akter et al. (2026) studied this question.