OpenAI’s Colin Jarvis says enterprise AI is stuck on deployment, not models

OpenAI’s Colin Jarvis at HumanX in Amsterdam, 23 September 2026. Credit: ALX MEDIA / HumanX Most companies that struggle with enterprise AI are not waiting for better models, OpenAI’s head of forward deployed engineering said on Wednesday. In about 80% of cases, Colin Jarvis said, the problem is deployment. Companies do not yet know how to roll AI out, govern it or prove they can trust it.

Glasgow-based Jarvis leads OpenAI’s forward deployed engineers (FDEs), who work inside customer companies to get its models into production. He spoke to Alex Hern of The Economist at HumanX in Amsterdam. “I don’t personally feel a lot of pressure to race ahead on model development itself,” Jarvis said. In the other 20% of cases, customers say the model cannot yet do the job, he said. Those gaps are now narrow and specialist, such as a task in semiconductor design, rather than everyday office work.

What FDEs do The team started out as mostly software engineers, Jarvis said, because building on the raw API meant writing everything from scratch. With tools such as Codex, about 50% of each project is now custom work, down from about 90%. That has shifted hiring toward domain experts, including a former investment banker, former scientists and a chip verification engineer. Every engagement starts with a two-day visit.

The team asks business leaders to ignore AI and name the biggest levers in their business. It then works on whichever that turns out to be. At one semiconductor company, engineers spent maybe 80% of their time on bugs from overnight jobs, Jarvis said. OpenAI built a system that finds the root cause, proposes a patch and applies it after an engineer reviews it.

It spread from one department to the whole business. The company estimates it saves roughly $40m to $50m a year. Other projects include an agent that iterates car part designs from plain-language instructions, for a European manufacturer. Another helps draft clinical trial documents, where a human always stays in charge.

Why pilots fail Jarvis named two common mistakes. Companies pick a use case because it seems to fit AI, not because it matters. Or they get one good use case working in one department, and it stays a demo that never spreads. The companies that succeed measure success by production use, not proofs of concept, he said.

One semiconductor customer has about 35 use cases live after roughly 18 months. It built a central team to scale projects and placed small groups of engineers in each business unit. FDEs have no financial incentive tied to usage, Jarvis said. They are measured on whether a project reaches production and moves a real metric.

When OpenAI’s embeddings were too slow for a Klarna search service, he told the company to use an open-source model instead. “From OpenAI’s side, we should always be temporary,” he said. The pace question Hern noted that Sam Altman was at the UN that day, as pressure grows to slow AI down. Jarvis said OpenAI has shown it will pause when it reaches the limits of its safety frameworks. He said he thought it paused its main reinforcement learning run “in September this year”.

OpenAI published its own post on the pause on 18 August. It describes a two-week pause in reinforcement learning training on its latest models. It also says its largest planned frontier run stayed on hold. Altman has since said OpenAI will set its own pace without waiting for Congress.

Jarvis said FDEs help test whether safety frameworks built in a lab hold up in messy real companies. OpenAI is not alone in sending engineers into clients’ offices. AWS is spending $1bn on the same model, and Microsoft launched a $2.5bn deployment business in July.

Leave a Reply

Your email address will not be published. Required fields are marked *