ZECH
Guide · 7 min read

From AI pilot to production: a readiness checklist

A pilot is ready for production when it has a named owner, an evaluation set that reflects real work, integrations that run under their own identity, documented controls, and a plan for monitoring, cost and support. If any of those are missing, fix them before you scale.

A pilot proves that something is possible. Production proves that it keeps working on real inputs, for real users, at a cost the business accepts, while someone is accountable for it. The gap between the two is rarely the model. It is usually everything around it: data access, integration, evaluation, controls and day-two operations.

This checklist is the one we use before recommending that a pilot move forward. It is organized into six areas. Treat each item as a yes-or-no question; a "no" is not a failure, but it is work to schedule before launch.

The short answer

A pilot is ready for production when you can say yes to all of the following:

  • A named business owner is accountable for the outcome and has agreed on what success means.
  • There is an evaluation set drawn from real cases, and the system passes an agreed threshold on it.
  • The system reaches data and tools through service identities and approved interfaces, not a developer's credentials or a manual export.
  • Risks have been assessed, controls are documented and people know where human review happens.
  • Monitoring, cost tracking, incident response and a rollback path are in place.
  • The people who will use it have been involved, trained and given a way to report problems.

The sections below break each area down.

1. Readiness: is the problem worth scaling?

Pilots often succeed on a narrow slice of work chosen because it was convenient. Before scaling, check that the result generalizes to the work that actually matters.

  • Owner. One person in the business owns the outcome, not only the project.
  • Baseline. You know how long the task takes today, how often it goes wrong and what it costs. Without a baseline you cannot show improvement.
  • Scope. You have written down which cases the system handles and which it does not.
  • Volume. The volume in production justifies the ongoing cost of running and maintaining the system.
  • Users. The people whose work changes have seen the pilot and their feedback has been addressed.

2. Architecture: will it hold up?

A pilot built in a notebook or a single script usually needs restructuring. That is normal and should be planned, not discovered.

  • Separation of concerns. Prompts, retrieval configuration, model choice and business rules are separate and versioned, so one can change without breaking the others.
  • Model portability. Switching model provider or version is a configuration change followed by an evaluation run, not a rewrite.
  • Failure behavior. You have decided what happens when the model is slow, unavailable or returns something unusable. Usually that means a clear fallback to a human queue.
  • Latency and throughput. The system has been tested at expected peak volume, not only at pilot volume.
  • Environments. Development, test and production are separate, with production data handled according to its classification.

3. Evaluation: how will you know it works?

In our view, this is the area most often skipped and the one that causes the most pain later. Without evaluation, every change is a guess.

  • Evaluation set. A set of real inputs with known good outputs, including the awkward and ambiguous cases, not only the easy ones.
  • Metrics that match the task. Accuracy of extracted fields, correctness of an answer against its source, whether a classification was right. Pick measures a business reviewer would recognize.
  • Threshold. An agreed level of quality that must be met before release and maintained after it.
  • Automated runs. The evaluation runs automatically whenever prompts, models or retrieval sources change.
  • Human review sample. A regular sample of production outputs is reviewed by people who know the work, and their findings feed back into the evaluation set.

4. Integration: does it fit the way work happens?

A pilot that requires people to copy and paste between windows will not survive contact with a busy team.

  • Service identities. The system uses its own credentials with the minimum permissions it needs.
  • Permissions carried through. If a user cannot see a document in the source system, the AI system does not show it to them either.
  • Where the output lands. Results appear in the tool people already work in, whether a ticketing system, a CRM or a document workspace.
  • Write actions. Any action that changes a record has limits, is logged and can be reversed or reviewed.
  • Upstream changes. Someone will know when a connected system changes its interface or data format.

5. Governance: can you explain and defend it?

Governance does not need to be heavy, but it needs to exist before launch rather than after an incident.

  • Risk assessment. You have considered what a wrong output could cause and set review points accordingly.
  • Human in the loop. It is clear which outputs a person approves, which are sampled and which are fully automated.
  • Data handling. You know what data is sent to which model, where it is processed and whether it is retained.
  • Audit trail. Inputs, outputs, sources and model versions are logged so a decision can be reconstructed.
  • Transparency. Users know when they are working with AI output and what its limits are.

6. Operations: who looks after it on day two?

Production systems drift. Inputs change, models are updated and users find new ways to use the tool.

  • Monitoring. Quality, latency, error rates and usage are tracked on a dashboard someone actually checks.
  • Cost tracking. Cost per task or per user is visible, with alerts for unexpected increases.
  • Incident response. There is a documented way to disable the AI step and fall back to the manual process.
  • Release process. Changes are tested against the evaluation set and rolled out gradually where possible.
  • Support. Users know how to report a bad output, and those reports are triaged.
  • Ownership after launch. A team is responsible for the system once the project team moves on.

A practical order of work

If the checklist shows several gaps, a sensible order is:

  1. Establish the owner, baseline and evaluation set first. Everything else depends on being able to measure.
  2. Fix integration and permissions next, because they usually take the longest to agree.
  3. Document governance decisions while the architecture is being hardened.
  4. Put monitoring and incident response in place before the first real users arrive.
  5. Launch to a limited group, compare results against the baseline and expand in stages.

This is our opinion based on how projects tend to stall: teams that start with evaluation make faster, better-informed decisions at every later step.

Where to go next

For the operational side of running models in production, see MLOps & LLMOps. If you would like a second opinion on a pilot you already have, get in touch.