This study addresses the estimation of conditional probability tables in Bayesian networks. Traditional procedures encounter difficulties when the number of parent variables is large because the number of parameters grows exponentially. Logistic regression can mitigate this dimensionality challenge, but it imposes assumptions about how a variable depends on its parents. To extend its applicability to more general scenarios, we introduce a method for systematically constructing artificial variables derived from the original predictors. We develop a general theoretical framework to compare alternative formulations of these variables. Two main approaches are proposed. The first relies on XOR combinations of the original variables, whereas the second is based on Linear Discriminant Analysis and discretization. The proposed approach is validated through an extensive empirical evaluation on large- and medium-scale networks, where we benchmark logistic regression against established techniques, including Bayesian inference procedures and probability tree models. The experiments demonstrate that logistic regression and discretization yield excellent performance while requiring fewer parameters than a full conditional probability table. • In Bayesian networks, parameters grow exponentially with the number of parents. • Enhanced logistic regression uses synthetic variables to improve model flexibility. • New variables are created via XOR combinations and LDA-based discretization. • The paper compares transformations via completeness, equivalence, and subsumption. • LDA-based logistic regression excels with fewer parameters than traditional tables.
Moral et al. (Wed,) studied this question.