This study explores strategies to guide the generation of polyimides with high glass transition temperatures (Tg > 750 K) through reinforcement learning. We present a systematic computational framework for analyzing and combining multiple scoring functions into a single score in reinforcement learning (RL) for molecular design. Rather than relying solely on a single scoring function based on a predictive model, we examine a range of complementary scores, including a novel naı̈ve high-Tg score and various Tanimoto similarity-based scores. We analyze these scores both individually and in combination with the predictive model-based score in order to assess their influence on the structural diversity and quality of the generated polymers. In addition, we investigate several methods for combining scores, such as arithmetic, geometric, and harmonic means, as well as a novel exponential-logarithmic function, referred to as ExpAgg. We evaluate how these aggregation strategies affect the outcomes of molecular generation across different reinforcement learning configurations. Our findings show that the choice of score combination method significantly impacts both the quality and diversity of generated polymers. The proposed ExpAgg achieves superior performance in multiple settings, revealing nontrivial interactions between score compatibility and model convergence. While the predictive model exhibits underestimation in the out-of-distribution region (>800 K), our multiscore framework successfully generates chemically reasonable high-Tg candidates. Based on these insights, we provide practical guidelines for selecting aggregation functions when fusing two scores. This case study on high-Tg polyimide generation demonstrates how score aggregation strategies influence molecular RL outcomes; broader generalizability to other molecular design tasks remains to be investigated. This work emphasizes the importance of moving beyond simple weighted averages in order to enhance targeted molecular design.
Tchagoue et al. (Thu,) studied this question.