Find what will break before your customers do.
For teams with a working AI prototype, an agent, or an app built quickly with Cursor, Claude, Lovable or Replit, who need to know what to check before real users, real data or real decisions depend on it.
The seven parts below are the same seven parts of a real-world AI system from our homepage. Work through them in order. Tick what's done, expand any item to see what "done" looks like, and print or save the page. Your ticks stay in your browser and are never sent to us. Scoring happens in a review, not here.
This is a practical checklist, not a certification and not a substitute for security, privacy or legal advice.
0 of 42 done
01
Identity and permissions
Who can use the system, and what can each person reach through it?
-
What done looks like
No anonymous path reaches the model or the tools. Service accounts have their own identities, not a shared key.
-
What done looks like
Permissions are checked at retrieval time with the user's identity, not assumed from the prompt. A user cannot ask their way into someone else's records.
-
What done looks like
Prompt templates, tool lists and model settings are changed through a controlled path with an audit trail, not from the chat window.
-
What done looks like
API keys and connection strings live in the runtime's secret store. A search of the built front end and git history finds none.
02
Business information
What information may the system use, and is it the right information?
-
What done looks like
A written inventory: source, owner, sensitivity, how it is refreshed, and whether customer or employee personal information is in it.
-
What done looks like
Personal, confidential and public information are labelled, and the system's access to each tier is deliberate rather than "everything it can reach".
-
What done looks like
For a set of real questions, someone has looked at what was retrieved and confirmed it was relevant, current and complete enough to answer from.
-
What done looks like
When the system answers from an outdated document, there is an owner and a process to correct the source, not just the answer.
-
What done looks like
You know which providers receive prompts and documents, where they process them, how long they keep them, and whether they train on them. This is written down and matches your privacy commitments.
03
AI model or agent
Does the approach fit the job, and is its behaviour defined?
-
What done looks like
Someone has written why this is a model task rather than a rule, a search or a form, and what simpler approach was considered.
-
What done looks like
Prompt changes are reviewed and can be rolled back. You can say which version answered a given request.
-
What done looks like
The system declines, asks or escalates in ways you have specified and tested, rather than improvising.
-
What done looks like
The model version is pinned. Upgrades go through the evaluation set before reaching users.
-
What done looks like
Each run records the steps taken, the tools called and their inputs and outputs, so a person can follow what happened.
04
Approved tools and actions
What can the system do in the world, and what stops it doing more?
-
What done looks like
Each tool exists because a workflow needs it. Nothing is exposed "in case".
-
What done looks like
A tool checks the acting user's rights itself. The model's request is not treated as authorisation.
-
What done looks like
Sending, paying, deleting, publishing and changing records go through a confirmation step or a human approval, with the rule written down.
-
What done looks like
Amounts, recipients, record identifiers and query scopes are checked against limits before anything runs.
-
What done looks like
No tool grants permissions, creates credentials or changes the tool list. Untrusted content (documents, web pages, emails) cannot inject instructions that reach a tool unchecked.
05
Validation
How do you catch poor or unsafe results before they cause harm?
-
What done looks like
Structured outputs are validated against a schema. Free-text outputs that drive decisions are checked by rules, a second model pass, or a person, and you know which.
-
What done looks like
Dozens of real or realistic inputs with expected outcomes, including hard cases and cases the system should refuse. It runs before every release.
-
What done looks like
Hallucinated facts, wrong retrieval, partial answers, refusals, timeouts and tool errors each have a defined user-facing behaviour.
-
What done looks like
Tests show the system does not reveal other users' data, secrets, or internal instructions when asked directly or indirectly.
-
What done looks like
Uncertainty is surfaced to the user in words, not hidden behind a confident answer.
06
Monitoring and evaluation
How do you know it keeps working after launch?
-
What done looks like
Inputs, retrieved context, model version, tools called, outputs, latency and cost are recorded, with personal information handled per your retention rules.
-
What done looks like
A sample of real interactions is reviewed on a schedule, and the results are compared with the pre-launch evaluation.
-
What done looks like
A feedback control exists, flags reach a named owner, and recurring problems feed the evaluation set.
-
What done looks like
Error rate, latency, refusal rate, tool failures and cost per day have thresholds that page a person.
-
What done looks like
You can see which workflows and which accounts drive spend, and you would notice a runaway loop within the hour.
07
Human approval or fallback
When should a person take over, and does that path work?
-
What done looks like
Every user-facing flow has a route to human help that does not depend on the model working.
-
What done looks like
A short written list of decisions the system may only draft or suggest, never make. The interface reflects it.
-
What done looks like
When the model, retrieval or a tool is down, users get a clear message and a working alternative, not a spinner or an invented answer.
-
What done looks like
Escalated cases go to a queue a named person watches, with a response time you have agreed.
08
Cost and performance
Will it stay fast and affordable as use grows?
-
What done looks like
You know the typical and worst-case response time for real requests, and the interface handles the slow ones honestly.
-
What done looks like
A budget per user or per day exists, with rate limits and caps that stop runaway spend.
-
What done looks like
Repeated work is cached where safe, retries are bounded, and a failure cannot multiply cost.
-
What done looks like
The system has been run at a multiple of expected concurrency, and you know where it degrades.
09
Release and ownership
Can you ship, roll back and look after it?
-
What done looks like
A release is a versioned artefact, deployed the same way each time, and can be rolled back in minutes.
-
What done looks like
Authentication, authorisation, data handling and the tool layer have automated tests that run before release.
-
What done looks like
Someone has read the authentication, authorisation and data-access code the tools generated, and dependencies are pinned and checked for known issues.
-
What done looks like
One person is accountable after launch, with hours each month for monitoring, fixes and improvement, and a backup when they are away.
-
What done looks like
Where it is used, people are told an AI system is involved, what it can do, and how to get help.
Next step
Want this done with you, on your system?
A Northvian production readiness review works through these parts on your actual system and produces a scorecard, a risk register, architecture recommendations, a prioritized fix plan and a production roadmap.