A pilot proves that something is possible. Production proves that it keeps working on real inputs, for real users, at a cost the business accepts, while someone is accountable for it. The gap between the two is rarely the model. It is usually everything around it: data access, integration, evaluation, controls and day-two operations.
This checklist is the one we use before recommending that a pilot move forward. It is organized into six areas. Treat each item as a yes-or-no question; a "no" is not a failure, but it is work to schedule before launch.
The short answer
A pilot is ready for production when you can say yes to all of the following:
- A named business owner is accountable for the outcome and has agreed on what success means.
- There is an evaluation set drawn from real cases, and the system passes an agreed threshold on it.
- The system reaches data and tools through service identities and approved interfaces, not a developer's credentials or a manual export.
- Risks have been assessed, controls are documented and people know where human review happens.
- Monitoring, cost tracking, incident response and a rollback path are in place.
- The people who will use it have been involved, trained and given a way to report problems.
The sections below break each area down.
1. Readiness: is the problem worth scaling?
Pilots often succeed on a narrow slice of work chosen because it was convenient. Before scaling, check that the result generalizes to the work that actually matters.
- Owner. One person in the business owns the outcome, not only the project.
- Baseline. You know how long the task takes today, how often it goes wrong and what it costs. Without a baseline you cannot show improvement.
- Scope. You have written down which cases the system handles and which it does not.
- Volume. The volume in production justifies the ongoing cost of running and maintaining the system.
- Users. The people whose work changes have seen the pilot and their feedback has been addressed.
2. Architecture: will it hold up?
A pilot built in a notebook or a single script usually needs restructuring. That is normal and should be planned, not discovered.
- Separation of concerns. Prompts, retrieval configuration, model choice and business rules are separate and versioned, so one can change without breaking the others.
- Model portability. Switching model provider or version is a configuration change followed by an evaluation run, not a rewrite.
- Failure behavior. You have decided what happens when the model is slow, unavailable or returns something unusable. Usually that means a clear fallback to a human queue.
- Latency and throughput. The system has been tested at expected peak volume, not only at pilot volume.
- Environments. Development, test and production are separate, with production data handled according to its classification.
3. Evaluation: how will you know it works?
In our view, this is the area most often skipped and the one that causes the most pain later. Without evaluation, every change is a guess.
- Evaluation set. A set of real inputs with known good outputs, including the awkward and ambiguous cases, not only the easy ones.
- Metrics that match the task. Accuracy of extracted fields, correctness of an answer against its source, whether a classification was right. Pick measures a business reviewer would recognize.
- Threshold. An agreed level of quality that must be met before release and maintained after it.
- Automated runs. The evaluation runs automatically whenever prompts, models or retrieval sources change.
- Human review sample. A regular sample of production outputs is reviewed by people who know the work, and their findings feed back into the evaluation set.
4. Integration: does it fit the way work happens?
A pilot that requires people to copy and paste between windows will not survive contact with a busy team.
- Service identities. The system uses its own credentials with the minimum permissions it needs.
- Permissions carried through. If a user cannot see a document in the source system, the AI system does not show it to them either.
- Where the output lands. Results appear in the tool people already work in, whether a ticketing system, a CRM or a document workspace.
- Write actions. Any action that changes a record has limits, is logged and can be reversed or reviewed.
- Upstream changes. Someone will know when a connected system changes its interface or data format.
5. Governance: can you explain and defend it?
Governance does not need to be heavy, but it needs to exist before launch rather than after an incident.
- Risk assessment. You have considered what a wrong output could cause and set review points accordingly.
- Human in the loop. It is clear which outputs a person approves, which are sampled and which are fully automated.
- Data handling. You know what data is sent to which model, where it is processed and whether it is retained.
- Audit trail. Inputs, outputs, sources and model versions are logged so a decision can be reconstructed.
- Transparency. Users know when they are working with AI output and what its limits are.
6. Operations: who looks after it on day two?
Production systems drift. Inputs change, models are updated and users find new ways to use the tool.
- Monitoring. Quality, latency, error rates and usage are tracked on a dashboard someone actually checks.
- Cost tracking. Cost per task or per user is visible, with alerts for unexpected increases.
- Incident response. There is a documented way to disable the AI step and fall back to the manual process.
- Release process. Changes are tested against the evaluation set and rolled out gradually where possible.
- Support. Users know how to report a bad output, and those reports are triaged.
- Ownership after launch. A team is responsible for the system once the project team moves on.
A practical order of work
If the checklist shows several gaps, a sensible order is:
- Establish the owner, baseline and evaluation set first. Everything else depends on being able to measure.
- Fix integration and permissions next, because they usually take the longest to agree.
- Document governance decisions while the architecture is being hardened.
- Put monitoring and incident response in place before the first real users arrive.
- Launch to a limited group, compare results against the baseline and expand in stages.
This is our opinion based on how projects tend to stall: teams that start with evaluation make faster, better-informed decisions at every later step.
Where to go next
For the operational side of running models in production, see MLOps & LLMOps. If you would like a second opinion on a pilot you already have, get in touch.