AI/machine learning

From Pilot to Production: Understanding the Challenges of Scaling AI in Oil and Gas

As upstream operators move beyond isolated experiments, the hardest part of the AI journey is not building a model, it is making it stick.

Connected Smart Oil Field Concept with Pump Jacks at Sunset
Patterns are emerging among the operators who have crossed from pilot to production.
Source: onurdongel/Getty Images.

[Editor's Note: Mohamed Alzaabi is a member of the TWA Editorial Board and is the author of previous TWA articles.]

Walk the exhibition floor of almost any energy conference today and the message is the same: artificial intelligence (AI) is no longer a curiosity in oil and gas, it is a line item. Yet a more sobering data point has emerged from outside the industry that upstream leaders cannot afford to ignore. A 2025 study by MIT's NANDA initiative, based on more than 300 public AI deployments and over 150 executive interviews, found that 95% of generative AI pilots deliver no measurable impact on profit and loss. That figure is specific to generative AI, but the scaling barriers it exposes such as data readiness, organizational alignment, talent gaps, and change management apply with equal force to the machine-learning models and physics-informed optimization tools that dominate upstream applications.

Across AI types, only about one in 20 initiatives ever escapes pilot purgatory to create durable, enterprise-level value. As I wrote in a previous article for TWA, the industry has no shortage of headline successes, from Shell's seismic-shot reduction to ADNOC's autonomous gas-lift control. But these are the visible 5%. Behind them sits a much larger population of proofs of concept that impressed a steering committee, generated a glossy internal case study, and then quietly stalled. EY's oil and gas digital operations practice puts it plainly: evolving from a proof of concept on a small project to true scale is a complex undertaking, and it is the last mile that proves most daunting (EY, 2025).

Understanding why that last mile is so hard, and what separates operators who cross it from those who do not, is the focus of this article.

Scaling1.png
Fig. 1—Only a small fraction of enterprise AI pilots ever reach production with measurable business value.
Source: MIT NANDA, The GenAI Divide: State of AI in Business 2025.

Why Scaling Is a Different Problem From Piloting

A pilot is, by design, forgiving. It runs on a hand-picked data set, in a single field or platform, under the close supervision of the same engineers who built it. Production is the opposite: it must run continuously, on incomplete and messy real-world data, across dozens of assets with different vintages of instrumentation, and largely without the original development team in the room.

ARC Advisory Group's 2026 analysis of industrial AI in oil and gas frames this transition as fundamentally a governance and operating-model challenge rather than a technology one. The report identifies the same forces compressing every operator's margin for error: tightening capital discipline, an aging and thinning workforce, rising operational variability, and the steady convergence of IT and OT systems, each of which raises the stakes for getting AI deployment right the first time and every time after.

The Data Foundation Most Pilots Never Test

The most common point of failure is also the least glamorous: data. A pilot can tolerate a curated, cleaned data set assembled by hand. A production system cannot. It must draw continuously from seismic interpretations, well logs, SCADA streams, drilling parameters, and maintenance records that were never built to talk to one another.

EY notes that many operators have spent recent years simply standing up the data foundations necessary to properly deploy AI, work that is invisible but unavoidable. ARC goes further, arguing that an integrated, governed data layer, what it calls an Industrial Data Fabric, is the prerequisite for breaking down IT/OT silos so that AI models can be trusted with decisions that have safety and financial consequences. The Open Subsurface Data Universe initiative reflects an industrywide recognition of the same problem: without a standardized, harmonized data layer, every AI use case starts from scratch.

Organization Before Algorithm

Even with clean data, scaling fails when the organization around the model is not ready for it. EY's research distills this into three requirements that map closely to what we are seeing across the industry. First, leadership must visibly anchor on value, treating AI investment with the same rigor as any capital project rather than a string of disconnected experiments. Second, scaling AI is not like adding cloud capacity; it requires people, processes, data, and technology to grow together, which is why EY describes it as an effort that must scale exponentially across every one of those dimensions at once. Third, and most underestimated, is culture: engineers and operators who have trusted their own judgment for decades are understandably reluctant to hand decisions to a model they cannot interrogate. Without a genuine change-management effort built on communication and feedback, adoption stalls long before the technology does.

The MIT NANDA research reinforces this from a different angle. It found that the highest-performing organizations did not necessarily build the most sophisticated models; they partnered smartly, embedded tools into existing workflows rather than bolting them on, and resisted the temptation to chase every use case at once. The report's authors describe this as a "learning gap," not a model-quality gap. The lesson translates directly to upstream operations: a hybrid, physics-informed model that drilling engineers trust will scale further than a more elegant model that sits unused on a dashboard.

The Widening Skills Gap

Scaling AI also assumes an organization has people who can build, govern, and interrogate it, and that assumption is increasingly shaky. A 2025 EPAM survey of energy companies found that while 98% plan to hire AI-specific roles, 54% of respondents believe their own workforce lacks the skills to deploy generative AI effectively, and on average 40% of staff will need upskilling within 18 months just to keep pace.

This gap is widening at the same time the experienced workforce is shrinking. The latest Global Energy Talent Index survey, conducted by Airswift, found that professionals aged 45 and older now make up 48% of the traditional energy workforce, while only 19% are aged 25 to 34 (GETI/Airswift, 2026). Scaling AI without a deliberate plan to transfer institutional knowledge into the very systems meant to augment that knowledge is a race against the clock.

Scaling2.png
Fig. 2—Hiring ambitions for AI talent are running well ahead of workforce readiness.
Source: Empowering the Energy Workforce: Bridging AI Skill Gaps to Drive Transformation, EPAM.

Cybersecurity and the OT Tightrope

Every AI system that touches production data is, by definition, a new connection between IT and OT environments, and every new connection is a new attack surface. DNV's Energy Cyber Priority 2025 survey of 375 energy professionals found that 65% now rank cybersecurity as their leadership's single greatest business risk. That concern is well-founded: 71% report feeling more exposed to OT-specific cyber events than before, and 57% admit their OT defenses still lag behind IT (DNV Cyber, 2025). Scaling AI into control rooms and well sites without closing that gap does not just risk a failed deployment; in an industrial setting, it risks a safety incident. Operators who are getting this right are treating cybersecurity as a design constraint from day one by implementing network segmentation between IT and OT environments, deploying OT-specific monitoring tools, and embedding secure-by-design principles into AI architecture before a model ever touches a live asset.

What Separates the 5% From the Rest

Patterns are emerging among the operators who have crossed from pilot to production. They anchor each initiative to a specific, measurable business outcome rather than a technology showcase. They invest in data governance before model sophistication. They pair AI tools with the people who will use them daily, treating change management as a workstream with its own budget and timeline, not an afterthought. And they build hybrid systems, blending domain physics with machine learning, that engineers can audit and trust under pressure, the same pattern behind the production-grade successes at Shell and ADNOC. None of this is as exciting as the model itself. But as the industry's own experience with digital transformation over the past decade has shown, excitement was never the scarce resource. Disciplined execution was.

For Further Reading
Global Energy Talent Index, Airswift.
Navigating AI Strategies in Oil and Gas, ARC Advisory Group.
The GenAI Divide: State of AI in Business 2025 by A. Challapally, C. Pease, R. Raskar, et al., MIT NANDA.
Energy Cyber Priority 2025: Addressing Evolving Risks, Enabling Transformation, DNV.
Empowering the Energy Workforce: Bridging AI Skill Gaps to Drive Transformation, EPAM.
Scaling AI for Maximum Impact in Oil and Gas, EY.