More financial institutions now have something resembling an AI lab: an innovation cell, a data team with a mandate to experiment, a working group with its own budget. They produce pilots. The pilots work, and still, only a small share of them ends up running in production. The standard explanation is that the technology was not ready yet. In my view, that explanation is almost always wrong.
A pilot is built to answer a narrow question: is this technically feasible? With the tools available today, that question gets answered faster and with less effort every year. What determines whether something reaches production is a different set of questions, and none of them are about the model.
There is an easy misreading of this argument I want to rule out up front. I am not suggesting the pilot should carry the full apparatus of a production system. A pilot that demands complete governance, final integration and production-grade controls stops being a pilot, it becomes a slow project that can no longer fail cheaply, which was its main advantage in the first place.
The point is different. Not building something during the pilot is not the same as never having identified it. That is usually where labs diverge: the ones that produce systems, know from day one what is missing. Because the distance between a successful pilot and a system running inside an institution is not the last stretch of the same road. It is a change of category.
A pilot and a system are not judged by the same standard
A pilot is judged by what it teaches. A production system is judged by what it sustains: it runs every day, on real data, on decisions that affect customers and that someone will eventually have to explain. Those are different standards and combining them has a concrete cost: committees approving the move to production because "the pilot went well," without anyone changing the question.
Before starting, it is worth writing down what the pilot will decide. Not what it will demonstrate: what it will decide. A pilot that cannot go wrong is not a pilot; it is a demo. And a demo informs no decision; it only confirms an intuition someone already held.
That means agreeing in advance on what result would be bad enough to stop. It is uncomfortable, particularly when the lab still has to justify its own existence. But a lab that never discards anything stops being a lab and becomes a development team with a better name.
Pilot conditions are almost always borrowed conditions
Nearly every pilot runs on temporary arrangements nobody writes down: a data extract someone prepared by hand, once; an isolated environment with permissions granted as an exception; a business expert reviewing every output before anyone else sees it; a vendor operating under a trial agreement that does not contemplate production data.
All of that is legitimate. It is precisely what makes speed possible. The trouble starts when those conditions become invisible and the team sizes the production effort as if they never existed. The expert reviewing every output was a quality control. In production that control is either automated, formalised as a step in the process, or gone, and with it, much of the quality that made the pilot look good.
My recommendation is simple and cheap: keep a running list of borrowed conditions alongside the pilot. Every time the team solves something "for now," it goes on the list. By the end of the pilot, that list is the closest thing you will have to an honest estimate of what production costs and the challenges it will face.
The problem and the measure come first, not last
Before choosing a tool, the team should be able to say which specific decision changes, who makes it today, how long it takes and with what error rate. That is the baseline. Without one, any result looks good, because there is nothing to compare it against, and the conversation about value turns into an exchange of impressions.
The measure matters as much as the problem, and it should describe the process rather than the model. A solution can be right most of the time and still change nothing: if its output arrives in a format the analyst does not use, or after the cut-off when it was needed, accuracy is beside the point. The useful question is not how often it is right, but how much it shifts the time, cost or risk of the process it lives in.
One piece is not optional in a regulated institution: what counts as a tolerable error, and who defines it. That answer does not belong to the lab. It belongs to the process owner, together with risk and compliance. Asking at the outset costs one meeting. Asking at the committee that approves production can cost the project.
What the pilot does not build, but should still record
Four fronts are rarely resolved in a pilot and should nonetheless stay in view while the pilot runs.
Operating model
Who runs it once the lab is no longer watching. Who reviews the outputs, how often and against what criteria. Who authorises an exception. What happens when it fails on a Sunday. A solution without an operating model may reach production, only to find out the hard way what was missing.
Data
Where it comes from, under what permission it is used, and whether it can be reproduced automatically at the frequency the process requires. Pilot data often exists as a one-off extract, informally negotiated with whoever had access. That is not a data source. That is a favour.
Security and confidentiality
What information leaves the perimeter, to whom, under what contract, with what retention and with what commitment about its subsequent use. These are reasonable questions before the first file moves, and awkward ones afterwards.
Traceability
If someone has to explain tomorrow why the system said what it said on a given date, can it be reconstructed? That requires keeping versions, inputs, criteria and model from the outset. It is one of the few things that cannot be rebuilt after the fact: either it was recorded, or it does not exist.
None of this needs a formal framework. One page is enough: owner, data, security, evidence. If a field is still blank when the pilot ends, that is fine, the goal is not to fill it in, but to make it visible before someone signs off on production.
The jump from pilot to production does not get shorter by writing more documentation. It gets shorter by knowing, from day one, in which direction you have to jump. A lab that works this way will probably produce fewer pilots and more systems, and I would argue that is the metric that ends up mattering.
At Rhisco Group we work with organisations that need to bring technology into critical processes without losing control or traceability. If your team is about to decide which of this year’s pilots moves to production, or you want help with a pilot, let’s talk.
This article was co-created with the assistance of artificial intelligence under strict supervision, editing, and verification of our team.