AI Product Development

From an AI idea or prototype to a product people can actually use.

You have an idea, a demo that impresses, or a prototype that works for you and falls over for the second user. Northvian builds the product around it: the parts a demo skips, in the order that gets you to real users.

01

Is this you?

  • “I'm a founder with a prototype that works in the demo. People liked it. I don't know what it needs before customers can use it.”

  • “We want an internal copilot for our team: something that answers from our own documents and drafts in our style. We've tried ChatGPT and Copilot for this and they don't know us.”

    ChatGPTCopilot

  • “We have an existing product and customers are asking for an AI feature. We don't want to bolt a chat box on and call it done.”

  • “I built it in Lovable, Replit or Cursor over a few weekends. It works for me. It needs real accounts, a real backend and real data handling, and I've hit the edge of what I can do alone.”

    LovableReplitCursor

  • “Our team built an agent that impresses everyone in the room and fails quietly in ways we can't reproduce.”

02

What this work is

AI product development is building a product where an AI model does part of the job, and building everything else the product needs so people can rely on it. The model is one part. The rest is the same as any serious software, plus a few parts specific to AI: deciding what information the model may see, checking what it produces, and measuring whether it is still doing the job next month.

Technology follows the problem. We do not start with retrieval, agents or fine-tuning; we start with the user's job and add the technique that job needs. Retrieval, finding the right passages from your own information and handing them to the model before it answers, fits when the answers live in your documents. An agent, an AI system that takes steps and uses tools to complete a task rather than answering one question, fits when the path varies and a person would otherwise choose it. Fine-tuning, further training a model on your own examples so it behaves a particular way, fits rarely, when the behaviour cannot be got from instructions and examples, and it is the last thing we reach for. Most products need none of the three at first. A good prompt, structured outputs and a well-designed interface get further than people expect.

The parts a demo skips are the parts that decide whether the product lasts. Identity: who is this user? Authorization, the check of what this specific user is allowed to see and do, applied on every request and never assumed from the prompt. Data boundaries: which information the model may use for this user, and which providers see it. Evaluation, a set of real cases with known good answers, run before every change so you can tell whether the change helped or hurt. Monitoring: every request logged with enough to reconstruct it, and alerts on the things that break. Fallback: what the user sees when the model, the retrieval or a tool is down. Cost: what a typical interaction costs, and where the ceiling is.

Prototypes built with Claude, Cursor, Lovable or Replit are a good starting point, and we treat them as one. The tools are fast and the code is often fine. What they do not do is decide your data model, review the authentication they wrote, or tell you the API key is in the client bundle. We assess the prototype honestly: keep what is sound, harden what is close, rebuild what cannot carry real users. Usually it is a mix, and throw it away is rarer than people fear.

03

What we might tell you not to do

  • Don't add retrieval, agents or fine-tuning by default

    Each one adds cost, latency and a new way to fail. Add the one the user's job demands, once a plain approach has shown its limits on real cases.

  • Don't fine-tune to fix a prompt problem

    If the behaviour can be described in instructions and a handful of examples, fine-tuning is the expensive way to get a worse version of that. It is for behaviour you cannot instruct, and it needs an evaluation set before it starts.

  • Don't rebuild the prototype from scratch on principle

    The Lovable or Cursor build has taught you what users want. Keep the interaction and what is sound underneath it. Replace the authentication and the data layer if they were generated without anyone reading them.

  • Don't ship without an evaluation set

    Twenty real cases with expected outcomes catch more than any amount of clicking around, and they are the only way to know whether next month's model upgrade helped or hurt.

  • Don't let the chat box be the product

    A chat interface is the easiest thing to build and often the wrong shape for the job. If the user's task has a known structure, a form with AI inside it beats a blank box.

  • Don't put the model in front of every user on day one

    One user group, one workflow, propose-and-confirm where the output matters. Widen from the evidence, not the launch date.

04

What we actually do

  1. 01

    Discovery

    The user's job, the outcomes, what exists today (including the prototype), and what working means in numbers. You get a written scope: what is in, what is out, and the first release defined.

  2. 02

    Prototype assessment, when there is one

    Architecture, data handling, authentication, secrets, dependencies, and the generated code. You get a keep, harden or rebuild map, part by part.

  3. 03

    Design

    The user experience for the AI parts (what the user sees when the model is unsure, slow or wrong), the architecture, the data model, identity and authorization, model selection, and whether retrieval or an agent is warranted. You get an architecture note and the interface design.

  4. 04

    Build

    Backend, frontend, integrations and the model layer, with validation, logging and the evaluation set from the first version. You get working software in a limited release, not a demo.

  5. 05

    Validate

    The evaluation set on real cases, failure-case testing, a security review of the AI surfaces (including prompt injection: instructions hidden in content the model reads), load, and cost per interaction. You get a readiness report against the same parts as our production review.

  6. 06

    Deploy and operate

    Repeatable deployment, monitoring, alerts, a runbook, cost visibility, and an owner on your side. You get a product you can run and change without us, or with us by the month.

You can enter at any stage. A founder with an idea starts at discovery. A team with a Cursor build starts at the assessment. A product with an AI feature already live may start at validate.

05

Illustrative scenario

A demo that answered anything, and a product that had to answer correctly

Situation

We're a small product team at a company that sells compliance software to clinics. We built a demo assistant in a week: paste a regulation, ask a question, get a good answer. Sales showed it to customers and three asked when they could have it. Then we tried it on a real customer's document set. The answers got vague, sometimes wrong, and we couldn't tell why.

What the work looked like

Discovery pinned the job down: a clinic administrator asking whether a specific procedure is allowed under their province's rules, with an answer they can cite. That ruled out a general chat box. The design added retrieval over each customer's own regulation set, permissions per clinic, structured answers that show the cited passage, and an explicit not-covered response. We built an evaluation set from the questions customers had actually asked and compared approaches against it. A clinic could never see another clinic's material.

What changed

The product answered less, and answered correctly. Every response showed its source. The evaluation set runs before each release, and it caught a model upgrade that changed the refusal behaviour a month later. It launched to the three customers first, then more, with cost per clinic visible from the first day.

06

How an engagement runs

  1. 01

    Discover and scope

    The user's job, the outcomes, the first release defined, and what we will need.

    1 to 2 weeks
  2. 02

    Assess the prototype

    Keep, harden or rebuild, part by part, with the reasons.

    About 1 week, when there is one
  3. 03

    Design

    Interface for the AI parts, architecture, data model, identity and authorization, model selection.

    1 to 3 weeks
  4. 04

    Build to first release

    Working software in limited use, with validation, logging and the evaluation set from the start.

    6 to 12 weeks
  5. 05

    Validate and launch

    Evaluation, failure cases, security of the AI surfaces, load, cost, then a staged launch.

    2 to 4 weeks
  6. 06

    Operate

    Monitoring, evaluation on live traffic, improvements, model changes handled through the evaluation set.

    Ongoing, by the month

Formats

  • A fixed-scope build for a defined first release, quoted after discovery.
  • Engagements scoped monthly, for product work that evolves with its users.
  • A prototype assessment on its own, for teams that want the keep, harden or rebuild map before deciding who builds.
  • Advisory hours by the month, for a team doing the build themselves.

Bands are typical for a first release with one AI capability and a small number of integrations. More integrations, or data that needs cleaning first, extend them.

07

What you receive

  • Scope and first-release definition

    What is in, what is out, and what working means in numbers.

  • Prototype assessment

    Keep, harden or rebuild, part by part, with the reasons written down.

  • Architecture note

    How the parts fit, what the model may see and do, and why the approach was chosen over the alternatives.

  • Interface design for the AI parts

    What the user sees when the model is confident, unsure, slow or wrong.

  • The product

    Source in your repository, owned by you, with tests and a repeatable deployment.

  • Evaluation set and results

    Real cases with known good answers, the results before launch, and the harness to run it again.

  • Readiness report

    The same nine areas as our production readiness review, scored on your product before launch.

  • Monitoring and cost view

    Every request logged, alerts on what breaks, cost per feature and per user.

  • Runbook and handover

    What to do when it fails, how to change it, and who owns it on your side.

08

Technical and risk notes

What we look at

  • Identity and authorization: how users authenticate, how every data access checks the acting user's rights, and how tenants are kept apart.
  • Data boundaries: what the model may see per user and per tenant, and which providers receive what, under which terms.
  • The model layer: selection by task, structured outputs, prompt versioning, pinned model versions, behaviour on hostile or ambiguous input.
  • Retrieval, when used: how documents are split and indexed, permission filtering at query time, freshness, and whether the right passages come back.
  • Agent design, when used: the tool list, the permission each tool enforces itself, approval points, and a log of every step.
  • Evaluation: the set, how it runs before each release, who or what judges the answers, and the thresholds.
  • Observability: the ability to see what the system did and why, from logs, traces across model and tool calls, and metrics, including cost per feature and per user.
  • Security of the AI surfaces: prompt injection, how outputs are handled, secrets, rate limits, abuse.
  • Performance: latency budgets, streaming, caching, timeouts, and what the user sees when things are slow.
  • Delivery: repeatable deployment, rollback, and tests for the application code, not only the model.

What can go wrong

  • The prototype's authentication was generated and never read, and a user can reach another user's data by asking.
  • The API key is in the client bundle or the git history.
  • Retrieval returns the wrong passages and the model answers confidently from them.
  • There is no evaluation set, so a model upgrade changes behaviour and nobody notices until a customer does.
  • Content the model reads, a document, a web page, an email, carries instructions and the system follows them.
  • The chat box lets users ask for things the product was never designed to do, and it tries.
  • A runaway loop or an unbounded retry doubles the bill overnight.
  • Latency is fine for one user and unusable for ten at once.
  • There is no fallback, so a provider outage takes the product down with it.
  • Nobody on your side can deploy or change it after the build.

09

Questions people ask

What do you build, and what do you not build?

AI products for customers, internal copilots and knowledge assistants, AI features inside existing products, and agents where the work warrants one. We do not build general software with no AI in it (see Software Engineering for the exception), and we do not build things we think will not work; we say so at discovery instead.

We built a prototype with Claude, Cursor, Lovable or Replit. Do you keep it, rebuild it or harden it?

We assess it first, part by part. The interaction and the parts that are sound stay. The parts that were generated without review, usually authentication, authorization, data access and error handling, get hardened or replaced. A full rebuild is the exception, and when we recommend one we show you why.

What is the process, and where can we start?

Understand, design, build, validate, operate, the same five stages as everything we do, with a prototype assessment added when there is one. You can enter wherever you are: an idea starts at discovery, a Cursor build starts at the assessment, a live feature that worries you starts at validate.

How are engagements structured and priced?

A fixed scope for a defined first release, quoted after discovery, or an engagement scoped monthly when the product needs to evolve with its users. Assessments and advisory time are available on their own. In every format the code is in your repository and belongs to you.

Do you use RAG, agents or fine-tuning?

When the user's job needs them. Retrieval when the answers live in your documents; an agent when the path varies each time; fine-tuning rarely, and only for behaviour that cannot be got from instructions and examples. Most first releases need none of them, and we would rather prove that on your real cases than assume it either way.

Do you have case studies?

We do not publish client work, and we do not invent it. The scenario on this page is illustrative: a composite of the shape this work takes, not a client. If you want a reference, ask, and we will see what we can arrange.

Are validation and operation included, or extra?

Validation is part of every build: the evaluation set, failure cases, security of the AI surfaces, load and cost are in scope before we call a release done. Operation is either handed to your team with the runbook and monitoring, or done by us by the month. Both are decided at scoping, not discovered at launch.

Which models and providers do you use?

Whichever fits the task, the data terms you need, and the budget. We pin the version, put a provider change through the evaluation set before it reaches users, and design so a switch is possible. We are not tied to one vendor and we do not resell any.

Bring the idea, the demo or the Lovable build.

We'll tell you what it needs before real users, then build it with you.