DOI: 10.1190/int-2025-0046 ISSN: 2324-8858

Deep Reinforcement Learning for Strategic Drilling: Policy Optimisation on Seismic Data from Synthetic Geological Models

Roderick Perez Altamar

Abstract

Deep reinforcement learning autonomously derives profitable drilling policies directly from seismic data, reducing reliance on sparse, biased, and often inaccessible historical well records. By training a Deep Q-Network (DQN) agent within a custom economic simulator driven by Markov chain structural modelling, agent-based hydrocarbon migration, and acoustic fluid substitution, we demonstrate that an AI agent can learn early-stage exploration strategies entirely from synthetic geological environments. A key control on the learned drilling behavior is the strategic time horizon governed by the reinforcement learning discount factor (γ). A high discount factor (γ=0.99) strongly promotes deep, long-horizon drilling strategies by increasing the importance of delayed rewards and encouraging the agent to accept substantial upfront drilling costs in pursuit of larger future returns. However, the experiments also show that competitive cumulative rewards can be achieved by some low-discount-factor models when combined with favorable learning-rate and exploration schedules, indicating that ? interacts with the broader training configuration rather than acting as an isolated requirement for success. High γ nevertheless remains the clearest driver of sustained deep exploration, whereas low γ generally favors shallower and more immediately rewarding actions. Furthermore, the results reveal that exploration strategy and learning rate substantially modulate policy performance, with an inverted ? schedule combined with a high learning rate producing the highest cumulative reward in the tested configurations. When deployed on unseen, structurally complex geological environments, the optimized agent systematically targets features such as crestal anticlines and compartmentalized fault blocks, while also exhibiting economically adaptive behavior in less favorable scenarios. Overall, the results demonstrate that successful autonomous drilling strategy depends not only on geophysical pattern recognition but also on the joint mathematical encoding of long-term reward valuation, exploration behavior, and learning dynamics.

More from our Archive