The present article introduces TRANS-D3, an innovative hybrid method that combines the Twin Delayed Deep Deterministic Policy Gradient (TD3) reinforcement learning algorithm with the Transformer architecture for predicting the remaining useful life (RUL). The model utilizes an optimized reward function based on the Linear Quadratic Regulator (LQR) to approach error correction as a dynamic control problem. On the CMAPSS dataset, TRANS-D3 demonstrates a marked advantage, achieving RMSE reductions of 84–90% in baseline situations (FD001) and 23–45% in highly variable contexts (FD003/FD004). Statistical validation demonstrates high reliability, with a coefficient of determination R2 of more than 0.93 in each of the subsets; the maximum is 0.9984 in FD001. The 95% confidence intervals for the mean error, ranging from 0.709 to 1.244 in FD001 and from −1.324 to 1.748 in FD004, also confirm that the framework is a statistically unbiased estimator. In terms of Score, the model reduces penalties by between 80% and 95% compared to advanced architectures such as DAST or STAR, ensuring very stable predictions. These findings present a novel robust optimization paradigm, which is essential for ensuring the safety and reliability of complex industrial systems in the context of Industry 4.0.
Paredes et al. (Fri,) studied this question.