NeuryaBook your Discovery
Intelligent operations

Why AI pilots don't reach production

Rubén Galindo-Ávila10 min read

An AI pilot reaches production when it has three things: an owner accountable for the business result, a number with a baseline and a deadline, and a process redesigned to work with the agent. Without all three, it stays a demo.

If you have an AI investment on the table and you still haven't authorized it, there's one question worth answering before any other: what number are you going to measure it against?

Almost nobody answers it before signing. Something gets authorized that promises to improve a process without ever putting into numbers how that process stands today: what it costs, how long it takes, how much margin it burns. And what comes next was already described by Gartner in April 2026: at least 50% of generative AI projects are abandoned after the proof of concept, due to poor data quality, insufficient risk controls, escalating costs or unclear business value. None of the four causes is technological.

We didn't reach this conclusion by reading studies. The first company whose process we redesigned so an agent could work on it was our own, and that's where we learned the expensive part: technology is almost never what fails; what fails is the process you connected it to. That's why we don't lead with the technology, but with the process that moves your numbers.

This article answers a single question: what has to be in place, before you authorize, for a pilot to cross into production instead of staying a demo? It's three things. And the first is that number.

Before you authorize: what number you'll measure it against

The starting number is not a project deliverable. It's a requirement to authorize it. It's taken beforehand, on the process as it stands today, and it's the only figure everything that comes later will be able to be compared against.

You don't need a dashboard. You need three facts about the process the investment promises to improve:

  1. How much it moves per month: revenue that depends on it, or the direct cost of running it.
  2. How long it takes today: the full cycle, end to end, not the step that hurts most.
  3. How much margin stays trapped in it: rework, waits, corrections.

With those three you can do the math that almost never gets done: multiply what the process moves per month by the percentage the initiative promises to move, by the months it will take to show a signal. That's the value at stake, and also the cost of not deciding, because while the project doesn't cross, the process keeps leaving that money on the table.

Add what isn't on the books and is still real: the quarter of executive attention it consumes, and the fact that the next AI investment is already born under suspicion at the board.

A leadership team that can't say in money terms what will change isn't evaluating an investment: it's giving an opinion on it. That's the difference between deciding in advance and explaining the result afterward.

What exactly does «in production» mean?

An agent is in production when the process can no longer run without it. Not when it works, not when the team uses it "when they remember to": when removing it would force the flow to be rebuilt.

That's the line separating a pilot from a system. A pilot runs in parallel to the real process, which is why it's reversible without consequences, and also why it's invisible on the income statement. An agent in production is inside the flow, with permissions, with traceability and with someone accountable when it fails.

And how long does it take to put an AI agent into production? In our entry-level offering the declared scope is 90 days or less for the first agent on a real process. It's not a promise of a result: it's a scope limit. If a first case needs more than a quarter to show a signal, it's almost always because too big a process was chosen, not because the technology is slow.

Condition 1 · The process redesigned to work with the agent

The number-one reason a pilot doesn't cross is having asked a new tool to run on an old way of working. The pilot inherits the steps that existed only because there was no data before: the manual validation that made up for one system not talking to another, the report someone built because nobody trusted the previous one, the double data entry born from an error eight years ago.

Automating that doesn't fix it. It speeds it up.

We saw it clearly in the case we cite most: a leading bottler in Mexico with critical vacancies open for nearly two months, where every unfilled week meant a stopped line and overtime. The first option on the table was to automate the recruiting process exactly as it was. It was discarded. It would have produced the same flow, faster and more expensive to maintain. The whole flow was redesigned first and the agents came in afterward, over screening, evaluation and coordination.

That decision, not to automate what already existed, is what explains the result. And it's the one that generates the most resistance, because it forces you to touch the way of working before you've seen the benefit. That's why it's a condition of authorization and not an adjustment along the way: if the redesign isn't accepted before signing, it won't be accepted afterward.

Redesigning the process before automating it also means deciding what happens with what's left over and what's missing. There are steps that disappear because they no longer have a reason to exist once the agent takes over the work, and there are new steps that appear because the redesigned process needs them: exception review, rule tuning, agent supervision. Some roles that operate the process today stop making sense as they are defined after the implementation, and that conversation happens before signing, it isn't discovered afterward. The question that follows is never how many people are surplus: it's what to do with the time that's freed up and to which value-creating process that talent is redirected. That's why you never automate a person. You automate a process, and the person is assigned another one that does move the number.

The 2025 MIT study carries the most uncomfortable finding on this topic: the vast majority of enterprise generative AI pilots produced no measurable return on the income statement, and the root cause wasn't the quality of the models, but how integrated they were into the workflow. The ones that did create value were the ones that got inside the process, not beside it.

Condition 2 · A business owner accountable for the result

If the project owner is IT, a tool was approved, not a change in how money is made. And a tool doesn't have to move a number: that's not its job.

The owner of an AI case is whoever answers for the process to leadership. Collections, if the case is collections. The plant, if it's the plant. IT builds, enables and controls; it can't answer for a commercial result it doesn't govern.

This is where the governance and control of AI agents in production comes in, which is what a board asks before approving and almost nobody has answered: who can see what, what gets logged, which decisions the agent makes on its own and which require a human.

AI agents with a human in the loop is not a concession to fear: it's design. The human confirmation point is defined from the start, it isn't added after a scare: when it's tacked on at the end as a patch, the team uses it for two weeks and then skips it.

Condition 3 · The baseline, taken before you start

Without the number from before, there's no way to defend the number from after. It's the quietest of the three conditions, because it breaks nothing: it just leaves the project with no way to win.

A pilot without a baseline enters a comfortable limbo. It neither wins nor loses: it survives. Three months later nobody can say whether it improved, not because it didn't, but because there's nothing to compare it against. And reconstructing the baseline at the end doesn't work: the board knows how to tell a measured number from a remembered one.

How do you measure an AI agent's performance? With a single business metric, taken before you start and with a cutoff date on the calendar. One, not a dashboard. We put it this way in every case we deliver:

We don't report hours saved. We report recovered revenue, cycle time, cost per transaction and capacity gained without adding headcount, with a baseline taken before starting and a review with leadership at every phase.

That's also the starting criterion: a case without a baseline doesn't get started. It's the most uncomfortable conversation of the first week and the one that avoids the impossible argument of month six.

What looks different when it does cross

A case that crosses is recognized by its size, not its technology: small scope, big number, an owner with a name and a cutoff date.

At the bottler, the recruiting cycle went from 50 to 10 days (-80%), with the critical positions filled 40 days earlier. The figure could be stated because the baseline existed from day 0. The technology wasn't the hard part.

The three conditions hold each other up, and that's why they fail together. The business owner is the one who agrees to redesign the process, because they are the only one with the authority to change how their area works. The redesigned process is what actually makes the number move. And the baseline is what lets you prove it to whoever authorized the investment.

Take one away and the other two stop working: a redesigned process without an owner reverts in three months; an owner with a baseline but no redesign defends a number that didn't move; and a redesign with an owner but no baseline produces an improvement nobody can credit.

That's why the order matters. First where the money is leaking and against what baseline. Only after that, what gets automated and in what order.

What to do on Monday

  1. Take the starting number of the process you're about to authorize. How much it moves per month, how long it takes, how much margin stays trapped in it. If you can't write it in one line, that's the finding, and it's reason enough not to sign yet.
  2. Give the process owner a first and last name, not the project's. If the name is IT's, there's no owner yet.
  3. Mark in the current flow the steps that exist only because there was no data before. Those are the candidates to disappear, not to automate.
  4. Set the cutoff date: if by day X the number isn't there, it stops. Write it on the calendar and communicate it; if it only lives in the meeting notes, it doesn't exist.

Start from where you are

The useful question isn't whether AI works. It's which of the three conditions you're missing today: before authorizing, or in the case you already have running.

Our self-diagnostic is 14 questions and five minutes, and you come out with a read on where your company stands on AI: data maturity and business maturity separately, so you can see the imbalance.

If you already know which process it is and what's missing is the number, that's where an Agentic Discovery begins: diagnostic, business case and a first agent in production, with the estimated return on the table before you authorize.

To go deeper:

Documented cases by industry available under a confidentiality agreement.

Frequently asked questions

How long should an AI pilot last?

A first case should show a signal in 90 days or less. If it needs more than a quarter, the problem is almost always the scope: too big a process was chosen, or one with too many owners. Shrink the scope before extending the deadline.

Who should own an AI project?

Whoever answers for the process to leadership, not whoever builds it. If it is collections, the head of collections. IT enables, controls and maintains the system, but cannot answer for a business result it does not govern.

What's the difference between a pilot and a Quick Win?

A pilot tests whether the technology works. A Quick Win tests whether the business improves: it runs on a real process, with a baseline, a business owner and a cutoff date, and delivers a number of its own in 90 days or less.