• Develops the concept of geophysical control, where time-lapse geophysical monitoring informs decision-making in geological carbon storage (GCS). • Demonstrates that transformer-based models (DTQN, ODT) outperform conventional value-based agents (DQN, DRQN, DTQN) in learning efficiency, stability, and cumulative reward. • Shows that ODT achieves superior performance through a two-phase offline–online training strategy and explicit modeling of temporal dependencies. • t-SNE analysis reveals that incorporating sequential, multimodal data (seismic + gravity + control histories) helps mitigate partial observability. • Attention analysis shows ODT dynamically prioritizes seismic observations during critical periods, diverging from traditional monitoring expectations that emphasize late-stage sensitivity. • Results highlight that transformer-based generative DRL agents offer cost-effective, high-performance control strategies and provide new insights into the temporal value of geophysical data for proactive leakage prevention. Geophysical control uses time-lapse geophysical measurements to optimize decision-making in sustainable energy applications such as geological carbon storage and geothermal power generation. Deep reinforcement learning (DRL) provides a natural framework for such control, yet its real-world deployment is challenging due to the limited observability of the subsurface state. To address this problem, we hypothesize that utilizing DRL architectures capable of modeling temporal relationships in geophysical data can alleviate such partial observability issues. Specifically, we compare four DRL agent architectures: three value-based methods and one autoregressive generative policy model, the Online Decision Transformer (ODT), in the context of geological carbon storage optimization. Using time-lapse seismic and borehole gravity data as geophysical control measurements, the ODT outperforms value-based agents in both decision-making abilities and learning speed, even with uncertainties in the underlying geological model. We attribute such superior performance to a combination of a two-phase, offline-online training strategy, effective modeling of temporal dependencies using transformers, and the integration of multimodal inputs—including combined geophysical measurements (states), historical control parameters (actions), and conditioned feedback metrics (rewards or return-to-go). t-SNE analysis further demonstrates that incorporating sequential, multimodal data from multiphysical monitoring and control-feedback histories helps mitigate partial observability. Attention-based analysis reveals that the agent dynamically prioritizes geophysical observations—especially seismic measurements—during critical control periods while appropriately downweighting less informative signals. Interestingly, these patterns depart from traditional geophysical monitoring expectations (which typically emphasize higher sensitivity measurements at later time steps when leakage becomes more pronounced), and offer new insights into the temporal value of geophysical data for proactive decision-making in leakage prevention. These findings highlight the potential of transformer-based DRL agents for enabling cost-effective, high-performance geophysical control in realistic subsurface systems. Conceptual workflow of geophysical control of geological carbon storage (GCS) operations using transformer-based deep reinforcement learning. Numerical GCS simulations provide time-lapse geophysical measurements (seismic and borehole gravity), which are processed by the generative policy AI model of Online Decision Transformer (ODT). Through sequential modeling of multimodal data, past actions, and return-to-go, the ODT generates optimized control policies that maximize CO2 stored while minimizing leakage risk.
Noh et al. (Sun,) studied this question.