PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 12, 2026Concurrency and Computation Practice and Experience1 citations

An Adaptive Computation Model for Fault‐Tolerant Execution of Complex Scientific Workflows in Distributed and Federated Cloud Infrastructures

View Full Paper
VPV. PadmavathiRKR. Kanimozhi

Key Points

  • The aim is to develop an adaptive computation model for executing complex scientific workflows that addresses resource heterogeneity and dynamic failures.
  • Introduced a dynamic task migration engine for real-time task relocation.
  • Implemented a predictive fault detection module using time-series analysis.
  • Developed an adaptive resource provisioning mechanism to dynamically adjust cloud resource allocation.
  • Tested on a hybrid cloud environment with Apache Airflow and major cloud services.
  • Validated using benchmark scientific workflows in genomics and environmental modeling.
  • Achieved a 47% decrease in fault recovery time.
  • Increased workflow fulfillment speed by 28%.
  • Realized a 22% time savings compared to baseline methods.
  • Demonstrated model scalability and cloud-agnostic properties.

Abstract

ABSTRACT Complex scientific workflows that involve resource heterogeneity, dynamic failures, and changing workloads pose major challenges when executed across distributed, federated cloud infrastructures. The static execution models traditionally used lack the flexibility to maintain such environments at the required performance and reliability levels. The paper offers an adaptive composition model of fault‐tolerant execution of workflows. The model incorporates: (1) Dynamic Task Migration Engine, which facilitates real‐time task migration to new systems based on failures or performance degradation, (2) predictive fault detection module, which intrinsically predicts a potential system failure using time‐series analysis and system parameters, and (3) adaptive resource provisioning level which dynamically scales up and down the cloud resources according to computation demand. We have tested the model on a hybrid cloud testbed comprising Apache Airflow with custom extensions, as well as on AWS, Google Cloud, and OpenStack. Validations were conducted on benchmark scientific workflows in genomics and environmental modeling across a range of fault scenarios. The final findings indicate a 47% decrease in fault recovery time, a 28% increase in workflow fulfillment speed, and 22% time savings compared to the baseline approaches. The model is scalable and cloud‐agnostic, enabling resilient scientific computing, especially in dynamic and high‐throughput environments. The implications of future work include applying reinforcement learning to optimize policies and support edge‐cloud systems.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Padmavathi et al. (2026) studied this question.

synapsesocial.com/papers/698d6ebb5be6419ac0d54740https://doi.org/10.1002/cpe.70611
Ask AI
Helpful
Bookmark
Share
View Full Paper