ABSTRACT Complex scientific workflows that involve resource heterogeneity, dynamic failures, and changing workloads pose major challenges when executed across distributed, federated cloud infrastructures. The static execution models traditionally used lack the flexibility to maintain such environments at the required performance and reliability levels. The paper offers an adaptive composition model of fault‐tolerant execution of workflows. The model incorporates: (1) Dynamic Task Migration Engine, which facilitates real‐time task migration to new systems based on failures or performance degradation, (2) predictive fault detection module, which intrinsically predicts a potential system failure using time‐series analysis and system parameters, and (3) adaptive resource provisioning level which dynamically scales up and down the cloud resources according to computation demand. We have tested the model on a hybrid cloud testbed comprising Apache Airflow with custom extensions, as well as on AWS, Google Cloud, and OpenStack. Validations were conducted on benchmark scientific workflows in genomics and environmental modeling across a range of fault scenarios. The final findings indicate a 47% decrease in fault recovery time, a 28% increase in workflow fulfillment speed, and 22% time savings compared to the baseline approaches. The model is scalable and cloud‐agnostic, enabling resilient scientific computing, especially in dynamic and high‐throughput environments. The implications of future work include applying reinforcement learning to optimize policies and support edge‐cloud systems.
Padmavathi et al. (Sun,) studied this question.