AI Production Readiness

Find what will break before your customers do.

You built something quickly and it works. Now real users, real data or real decisions are about to depend on it, and you're not sure it's ready. Northvian reviews the whole system, not just the model, and tells you what to fix first.

01

Is this you?

  • “We built this quickly and it works. Is it actually ready?”

  • “Our AI feature gives good answers in the demo and inconsistent ones for real users, and we can't reproduce the bad ones.”

  • “I built the app with Cursor or Claude and I'm about to take payments. Nobody else has read the authentication code.”

    CursorClaude

  • “It's built in Lovable or Replit, it has real users now, and I'm not certain where the API keys live.”

    LovableReplit

  • “The API bill doubled last month and we don't know which feature did it.”

  • “We inherited an AI system from a contractor or a colleague who has left, and nobody can say what it does when it fails.”

02

What this work is

A production readiness review is a structured look at an AI system that is about to carry real users, real data or real decisions, or already does. It answers one question: what will break, and in what order should you fix it? It covers the model, and it covers the six other parts of a real system that a prototype never had to have.

"We built this quickly. Is it actually ready?" is the right question, and the honest answer is usually "parts of it". Speed is not the problem. Apps built with Cursor, Claude, Lovable or Replit are often well structured and the code is often fine; the tools are good at code. What they do not do is review the authentication they wrote, get authorization right (the check of which user may see which data, on every request), keep the API key out of the client, or tell you what happens when the model is down. Those are the gaps the review finds, and they are the same gaps that appear in systems built by hand under a deadline.

The review is organized around the seven parts of a real-world AI system, the same seven on our homepage and in the public checklist: identity and permissions; business information; the AI model or agent (an AI system that takes steps and uses tools to complete a task, not just answer a question); approved tools and actions; validation; monitoring and evaluation (evaluation being a set of real cases with known good answers, run before every change); and human approval or fallback. We add two cross-cutting areas, cost and performance, and release and ownership. Every finding lands in one of those nine, so the report reads the way the checklist does and your team uses one vocabulary.

It is not a penetration test, though it covers the security items specific to AI systems: prompt injection (instructions hidden in content the model reads, which an agent may then act on), data leaking between users, secrets in the wrong place, missing rate limits. It is not a line-by-line audit of every file. It is the review that tells a founder or a product lead whether to launch, what to fix first, and what can wait.

03

The seven parts we review

A prototype proves an interaction. A system people depend on also has to know who's asking, what information it may use, what it's allowed to do, and what happens when it's wrong.

Prototype

  1. User
  2. Model
  3. Answer

Real-world system

  1. 01 Identity & permissions Who can access what?
  2. 02 Business information What information may the system use?
  3. 03 AI model or agent What approach fits the job?
  4. 04 Approved tools & actions What actions can it take?
  5. 05 Validation How do we catch poor results?
  6. 06 Monitoring & evaluation How do we know it keeps working?
  7. 07 Human approval or fallback When should someone take over?
Every finding in the report maps to one of these seven parts, plus cost and performance, and release and ownership.

04

What we actually do

  1. 01

    Scope

    What the system does, who depends on it, and what is about to change: a launch, payments, a new customer, more load. You get a written scope and a list of what we will need access to.

  2. 02

    Read the system

    Architecture, code (with attention to the parts AI tools generate: authentication, data access, secrets, error handling), prompts, tool definitions, data flows, infrastructure, dependencies. You get first findings within days, not at the end.

  3. 03

    Test the behaviour

    Real inputs, edge cases, hostile inputs, the failure cases (model down, tool errors, timeouts), isolation tests between users, and prompt injection through content the system reads. You get a reproducible case for every finding.

  4. 04

    Measure

    Latency under load, cost per interaction and per user, and the results of your evaluation set; where none exists, we build a starter one from your real traffic. You get numbers instead of impressions.

  5. 05

    Score and rank

    Each of the nine areas scored, every risk rated by likelihood and consequence, fixes ordered by what blocks a launch, what should follow, and what can wait. You get the scorecard, the risk register and the fix plan.

  6. 06

    Readout

    A working session with your team on what to fix, how, and in what order. You get the production roadmap and, if you want it, a scope for the fixes.

05

What you receive

  • Production Readiness Scorecard

    The nine areas, each rated, with the evidence behind each rating.

  • Risk Register

    Every finding, with likelihood, consequence, and the case that reproduces it.

  • Architecture Recommendations

    What to change in the design and why, in plain language for the owner and in detail for the engineer.

  • Prioritized Fix Plan

    Ordered: blocks launch, fix soon, improve later, with an effort band on each.

  • Production Roadmap

    The sequence from where you are to running it with confidence, including monitoring, evaluation and ownership.

06

What we might tell you not to do

  • Don't launch on the strength of the demo

    The demo proved the interaction. It did not prove what happens with a real customer's messy documents, a vague question, a hostile input, or two hundred users at once.

  • Don't buy a penetration test first

    A pentest finds network and application security holes, and it is worth having later. It will not tell you that retrieval (the step that finds the right passages in your documents before the model answers) returns the wrong ones, that there is no evaluation set, or that the agent can be talked into emailing the wrong person.

  • Don't rewrite the whole thing because it was AI-generated

    Most of it is probably fine. Fix the authentication, the data access, the secrets and the failure handling, which is where the review finds the problems, and keep the rest.

  • Don't swap the model first

    When outputs are inconsistent, the cause is usually retrieval, missing validation or the absence of an evaluation set. Changing the model without those is guessing.

  • Don't add monitoring after launch

    The first week of real traffic is the most valuable data you will ever get about the system. Log it from the first request.

  • Don't let the bill be your first alert

    Cost caps, per-user limits and a daily spend alert are an afternoon of work, and they stop the overnight surprise.

07

Illustrative scenario

A support assistant that was right in the demo and unpredictable in the queue

Situation

We're a software company with a support assistant that drafts replies from our help centre. In testing it was excellent. In production, the support team said it was sometimes brilliant and sometimes made things up, and nobody could say which tickets went wrong or why. We were about to swap the model.

What the work looked like

The review logged a week of real tickets end to end and found the cause wasn't the model. Retrieval was pulling from the whole help centre, including archived articles for an old version of the product, and the prompt had no instruction to say when the sources didn't cover the question. There was no evaluation set, so every change had been judged by feel. Data isolation was fine. Cost per ticket was higher than it needed to be because the retrieved context was far larger than necessary.

What changed

Fixes in order: filter retrieval to current articles, add an explicit not-covered response, cut the context size, and build an evaluation set from real tickets the support lead had rated. Validation and monitoring went from red to green after a few weeks of the team's own work. The model stayed the same.

08

Illustrative scenario

A Cursor-built app, a week before its first paying customers

Situation

I'm a solo founder. I built a document-review tool for small accounting firms with Cursor and Claude over a few months. It works, I have a waitlist of firms, and I was about to switch on payments. A friend who runs engineering somewhere asked who had reviewed the auth. Nobody had.

What the work looked like

The review read the generated code with attention to the parts the tools wrote unsupervised. Authentication was sound. Authorization wasn't: any signed-in user could fetch another firm's documents by changing an id in the request. The model API key was in the client bundle. There were no rate limits, so one user could run up the bill for everyone. A model timeout showed a spinner forever. Uploads went to a provider whose terms allowed training on the content, which the privacy page said wouldn't happen.

What changed

Four items blocked launch and were fixed in a week: per-firm authorization checks on every data path, the key moved server-side, rate limits and a daily cap, and a provider setting changed with the privacy page corrected. Monitoring and a starter evaluation set followed. The app launched to the waitlist with a fix plan for the rest.

09

How an engagement runs

  1. 01

    Scope

    The call, access arranged, and the list of what we will review.

    Under 1 week
  2. 02

    Review and test

    Reading, behaviour testing, measurement, with first findings shared as we go.

    2 to 3 weeks
  3. 03

    Report and readout

    Scorecard, risk register, recommendations, fix plan, roadmap, and a working session.

    About 1 week
  4. 04

    Fix, if you want us to

    Scoped from the fix plan, with us or alongside your team, then a re-check of the scorecard.

    2 to 6 weeks

Formats

  • A fixed-scope production readiness review, quoted after the scope call.
  • A shorter pre-launch check for a single-purpose app with one AI capability.
  • Fix work scoped from the plan, with us or alongside your team.
  • A re-review after fixes, so the scorecard reflects the system that ships.

Bands assume timely access to code, infrastructure and a person who knows the system. A larger system with several agents or integrations sits at the upper end.

10

Technical and risk notes

What we look at

  • Product and use-case fit, and architecture: what the system is for, who depends on it, whether the AI part does a job users actually have, and how the parts fit.
  • Identity and authorization: authentication, the check of what this specific user may see and do on every data path, per-tenant isolation, admin surfaces, secrets.
  • Data boundaries, retrieval and data architecture: data sources, sensitivity, privacy, retrieval quality and freshness, and what leaves your environment under which terms.
  • Model selection, prompt and system behaviour: the model chosen for the task, prompt versioning, pinned model versions, behaviour on hostile or ambiguous input, and visible steps for agents.
  • Tool permissions and actions: the tool list, permissions enforced per tool, approval on consequential actions, bounded inputs, no path to wider access.
  • Validation, evaluation, hallucination and failure cases: output checks, the evaluation set, known failure modes and their handling, tests for leakage between users.
  • Security and privacy specific to AI: prompt injection through user and retrieved content, output handling, rate limits, abuse, and dependency and supply-chain risk in generated code.
  • Observability and logging, the ability to see what the system did and why: logs with full context, traces across model and tool calls, alerts, a user feedback path.
  • Fallback behaviour and human oversight: what happens when the model, retrieval or a tool is down or slow, and the path to a person.
  • Cost, latency, scaling and release: token and API cost per interaction and per user, infrastructure cost, caps, caching, retries, latency under load, testing, repeatable deployment, rollback, operational ownership.

What can go wrong

  • Authorization written by the tool checks that a user is signed in, not that they may see this record.
  • Secrets in the client bundle, the prompt, or the git history.
  • Retrieval sees everything, so the system answers one user with another user's data.
  • No evaluation set, so a model or prompt change ships on feel.
  • Instructions in a document, a web page or an email reach a tool unchecked.
  • A retry loop, or an agent stuck in a cycle, multiplies cost overnight.
  • The model provider has an outage and the product shows a spinner.
  • Logs record the answer but not the retrieved context or the model version, so failures cannot be reconstructed.
  • A vendor's data terms do not match what your privacy page promises.
  • One person knows how to deploy it.

11

Questions people ask

What does the review cover?

The seven parts of a real system: identity and permissions, business information, the model or agent, approved tools and actions, validation, monitoring and evaluation, and human approval or fallback. Plus cost and performance, and release and ownership. Security items specific to AI systems, prompt injection, data leakage, secrets and rate limits among them, are inside those parts, not a separate track.

What do we receive?

A Production Readiness Scorecard, a Risk Register with a reproducible case per finding, Architecture Recommendations, a Prioritized Fix Plan and a Production Roadmap, plus a working session to go through them. If you want the fixes done, we scope that from the plan, with us or alongside your team.

How long does it take, and how is it priced?

A typical review runs three to five weeks from scope call to readout, with first findings shared in the first week. It is priced as a fixed scope, quoted after the scope call, because the size of the system sets the size of the review. A single-purpose app with one AI capability takes the shorter pre-launch check.

How is this different from a penetration test or a code audit?

A penetration test looks for security holes in the network and the application and is worth having later. A code audit reads every file. This review asks whether the system will hold up with real users: the AI-specific security items, but also reliability, retrieval, evaluation, monitoring, fallback and cost, which neither of the others covers. Most teams need this first.

What happens after the report?

You fix in the order the plan gives, with your team or with ours. Either way we are available for the questions that come up, and a re-review after the fixes gives you a scorecard for the system that actually ships. The monitoring and evaluation items can stay with us by the month.

Who is it for?

Founders about to launch or take payments. Product teams with an AI feature that is live and worrying. Teams that inherited an AI-coded app from a contractor, a colleague or a weekend, and need to know what they have.

Our AI feature gives inconsistent answers. Is that a model problem?

Usually not. Real users bring real inputs, edge cases and load, and the cause is most often retrieval returning the wrong passages, missing validation, or no evaluation set to tell good from bad. The review finds which, with reproducible cases. Swapping the model before that is guessing.

We built it with Cursor, Claude, Lovable or Replit. Will you tell us to rewrite it?

Almost certainly not. The code these tools write is often fine; what they skip is review of authentication, authorization, data access, secrets and failure handling. The review reads exactly those parts and tells you which to fix and which to keep. Fast code is not the problem.

Bring the system. We'll tell you what breaks first.

Start with the public checklist if you'd rather look yourself. A review does it with you, on your actual system.