ob / ouahabi-benhenni FR

Article ·

From proof of concept to production: what really changes in an AI project

An AI demo takes a few days to build. A system that is still running six months later takes something else. Here is what I learned leading about ten AI projects to production.

Why do so many AI projects stop at the demo?

Because a proof of concept answers a very different question from production. The PoC asks: "can the model do it?". Production asks: "can the organisation rely on it, every day, at an acceptable cost, and evolve it?".

The first question is solved with a good model and a clean data sample. The second involves everything else: real data, users, access rights, compliance, infrastructure, budget and, above all, the team that will maintain the system. When a project stalls after a successful demo, it is almost always on one of these points, rarely on the model.

What changes between a PoC and a production system?

  • Data. The demo sample was hand-picked. Production brings crooked scans, multilingual documents, empty fields, unexpected formats.
  • Edge cases. A model can be confidently wrong. In a demo nobody notices; in production someone makes a decision based on that error.
  • Cost. One call to a large model is cheap; a hundred thousand a month is not. Inference becomes a budget line.
  • Accountability. You must be able to explain where an answer came from, who validated it, and what happened when it was wrong.
  • Time. Models, data and needs change. A system that cannot evolve becomes debt.

How do you scope an AI project so it reaches production?

By treating production as a requirement from day one, not as a final step. Concretely, before writing a line of code, I try to establish four things:

  1. The business success criterion. Not "95% accuracy", but "processing a file drops from two hours to twenty minutes, with human validation".
  2. The constraints. Budget, compliance, available infrastructure, data confidentiality. They are not obstacles: they are the inputs of the architecture.
  3. The minimal useful scope. The smallest system that brings real value to real users.
  4. The owner. Who will operate the system after the project, and what they need to know to do it.

This scoping usually takes little time, and it avoids months of work on a system nobody will be able to put into service.

What architecture for an AI system that lasts?

The principle I apply in all my projects: AI inside the system, not a system around the AI. The model is a component, with an input and output contract like any other. Around it, deterministic stages prepare the data and check the results.

In a multilingual document extraction system, for instance, several narrowly scoped agents share the work — reading, classifying, extracting — and a deterministic checking stage verifies consistency before the data is delivered. Each agent can be tested and fixed separately. It is less spectacular than one model that "does everything", but it is what makes the system reliable.

Same logic in a precision medicine pipeline: every stage, from variant analysis to the report, produces an inspectable intermediate result. The language model writes the final report; it does not decide. The decision stays human.

Who will maintain the system?

This is the most frequently forgotten question, and the one that sinks the most projects. An AI system handed to a team that does not understand it is living on borrowed time.

Three practices make the difference:

  • Write architecture decisions down with their reason: why this model, why this breakdown, which alternatives were rejected.
  • Involve the team from the design stage. On my projects, much of the team is junior: involving them early is what makes them autonomous at the end.
  • Plan the handover as a deliverable in its own right, with documentation and skills transfer.

A checklist before saying "it's in production"

  • The business success criterion is measured, not just estimated.
  • Model errors are caught by a checking stage or by a human.
  • Every result can be traced back to its input data.
  • The operating cost is known and accepted.
  • Data access rights are respected end to end.
  • Someone on the team, other than the author, knows how to evolve the system.

In short

Moving from PoC to production is not about a better model. It is about scoping, architecture and team. A successful AI transformation chains small reliable systems rather than one spectacular big project — and it is judged six months after the demo, when the system is still running and the team knows why.

Go further: my approach to AI transformation and to AI systems architecture.

A transformation project, an architecture to design?

Companies, startups, institutions: describe your context in a few lines. I answer personally, with a first opinion on feasibility and approach.

Propose an engagement LinkedIn ↗ GitHub ↗

Engagements in Algeria, France, Europe, North Africa and remote · Arabic, French, English · contact@ouahabi-benhenni.com