AI Automation & Agents
Automate useful work, not impressive demos.
There's work your people repeat every day, and you want to know what can be automated safely. Northvian maps the workflow, decides which kind of automation each step needs, and builds it with the judgment left to people and the boundaries in from the start.
01
Is this you?
-
“Requests arrive by email and move between people by forwarding. Nobody can say where a given one is right now.”
Shared inboxes
-
“Invoices, intake forms and contracts arrive as PDFs, and someone types the same fields into a spreadsheet, then into the system.”
PDFsSpreadsheets
-
“Somebody built a Zapier or Power Automate flow last year. It mostly works, and now nobody trusts it or knows how to change it.”
ZapierPower Automate
-
“New staff ask the same questions about our own procedures, and the answers are in SharePoint somewhere.”
SharePoint
-
“We've been told an agent could do all of this. We'd like a second opinion before we give software access to our systems.”
02
Three kinds of automation, and what each needs
Not every problem needs an agent. Most real workflows sit in the first two columns, and that is a good thing: it is cheaper, faster and easier to trust.
- 01
Traditional automation
A fixed sequence of steps with fixed rules. When the input is structured and the rules are known, this is the cheapest and most reliable option, and often all you need.
- Structured inputs: forms, fields, files with known layouts
- Rules someone can write down
- Connections to the systems involved
- A place for exceptions to land
- 02
AI-assisted workflow
The same fixed sequence, with an AI step where the input is unstructured: reading an email, extracting fields from a PDF, classifying a request. The steps around it stay deterministic.
- Real examples to test the AI step against
- A check on what the AI step produces before it is used
- A confidence threshold, and a path to a person below it
- Accuracy measured over time, not once
- 03
Agentic workflow
An AI agent is software that can understand a task, use approved information or tools, and take a sequence of actions toward an outcome. Given a goal, it decides which steps to take and uses tools to take them. Worth it when the path varies each time and a person would otherwise have to choose it. Costs more in supervision.
- A short, explicit list of tools it may call
- Permissions it cannot widen on its own
- Approval on any action with consequences
- A log of every step and every tool call
- A person it can hand the case to
We do not treat the right-hand column as the goal. The right column is the one the work needs.
03
What this work is
AI automation is taking a workflow your people repeat and building a system that does the predictable parts, with AI used only where rules cannot handle the input. What separates it from a demo is everything around the model: what the system may read, what it may do, who checks it, and what happens when it is wrong.
An agent is the most talked-about kind of automation and the least often needed. Most workflows are a fixed sequence with one or two unstructured inputs, and a deterministic flow with an AI step for the reading is cheaper, faster and easier to trust than an agent. Where the path genuinely varies each time, an agent earns its place, and then the guardrails, the limits on what it may read, do and spend, enforced outside the model rather than requested of it, matter more than the model does.
Safely means specific things. Permissions: the system acts as its own named identity with the least access the workflow needs. Approvals: sending, paying, deleting, changing records and anything that leaves the building wait for a person, unless you have decided otherwise in writing. Escalation: below a confidence threshold, or on anything unusual, the case goes to a queue a person watches. Failure handling: timeouts, tool errors and bad extractions each have a defined behaviour, not a silent retry. Monitoring: every run is logged so someone can see what happened. Data boundaries: you know which providers see which information, and under what terms.
We build on what you already run where it is sound. If your Zapier or Power Automate flow does the right thing, the job may be to document it, add validation and give it an owner. If it does the wrong thing in three places, the job is to rebuild those parts. Custom code comes in when the tools cannot express the rules, or when the workflow needs permissions, logging and tests the low-code tools do not give you. The internal knowledge workflows, answering staff questions from your own procedures, usually want retrieval: a system that finds the right passages in your documents and hands them to the model before it answers, with each person's permissions carried through.
04
What we actually do
- 01
Map the workflow as it runs
With the people who do it, exceptions included. You get a workflow map that separates the predictable steps from the judgment steps.
- 02
Decide the kind, step by step
Traditional, AI-assisted or agent, per step, with the reasons written down. You get a design note you can challenge before anything is built.
- 03
Set the boundaries
Identity, permissions, which tools, which actions need approval, where exceptions land, which information may leave. You get the boundary sheet before any code.
- 04
Test the AI step on real examples
Extraction, classification or drafting, run against an evaluation set: real cases from your own work with the right answer known in advance. You get an accuracy figure on your data, and the failure cases listed.
- 05
Build and integrate
On your tools where they fit, custom where they do not, with validation and logging from the first version. You get a working system in a limited pilot.
- 06
Run it beside the people
A pilot period where the system proposes and a person confirms, until the numbers say otherwise. You get the evidence to widen it, or the evidence not to.
- 07
Hand over
Documentation, a named owner, a runbook for failures, a monitoring view and a maintenance plan. You get something your team can run.
05
What we might tell you not to do
-
Don't automate a workflow nobody has mapped
If the people who do it cannot describe it, exceptions included, the automation will encode one person's guess. The map is the cheapest part of the work and the part most often skipped.
-
Don't give an agent write access on day one
Read-only and propose-only first. Let it draft the email, prepare the record change, stage the payment, and let a person click. Widen its permissions when the log shows it has earned them, and only for the actions the log covers.
-
Don't automate the judgment step first
The step that needs an experienced person, deciding whether a claim is valid or a customer's request is within terms, is the one to leave alone. The re-typing, the routing and the chasing around it are where the time is.
-
Don't let the model be the permission check
A tool the agent calls must check the acting user's rights itself. The model asking for something is not authorization, the check of what this specific user or system is allowed to do.
-
Don't measure success in hours saved
Measure minutes per case, error rate, and how many exceptions reach a person. Hours saved that nobody reclaims are not savings.
-
Don't rebuild a working Zapier flow because it is Zapier
If it is right, document it and give it an owner. Rebuild the parts that are wrong, and only those.
06
Illustrative scenario
An intake queue that lived in one person's inbox
Situation
We're a benefits brokerage. Client change requests, a new employee, a salary change, a termination, come in by email to a shared inbox in every format you can imagine. One coordinator reads each one, keys it into the carrier portal and replies. When she's away, the queue stops. We'd asked whether an agent could take it over.
What the work looked like
The map showed a predictable sequence with one unstructured input: the email. An AI step read each message and proposed the fields. Validation checked them against the client's plan rules. Anything below the confidence threshold, and every termination, went to the coordinator. Writes to the carrier portal stayed propose-and-confirm for the pilot. Every run was logged with the email, the extracted fields, the confidence and who confirmed.
What changed
No agent, in the end. A deterministic flow with an AI reader. The coordinator went from keying every request to confirming most of them in a click, and the exception queue became visible to her manager for the first time. After the pilot, the routine change types moved to auto-confirm. Terminations still wait for a person, by design.
07
How an engagement runs
- 01 1 to 2 weeks
Map and design
The workflow map, the kind decision per step, the boundary sheet.
- 02 1 to 2 weeks
Prove the AI step
An evaluation set from real cases, accuracy on your data, the failure list.
- 03 3 to 6 weeks
Build and pilot
A working system in limited use, propose-and-confirm, monitoring from the first run.
- 04 1 to 2 weeks
Widen and hand over
Permissions widened where the log supports it, runbook, owner, maintenance plan.
Formats
- A fixed-scope workflow build, from map to hand-over, quoted after the map.
- A workflow assessment on its own, for teams that want the design and the evidence before committing to a build.
- Ongoing maintenance and improvement by the month, once something is live.
Bands are typical for one workflow. Systems with no API, or data that needs cleaning first, extend them, and we say so at the map stage.
08
What you receive
-
Workflow map
The steps as they run today, exceptions included, predictable and judgment steps marked.
-
Design note
Which kind of automation each step gets, and why. Written to be challenged.
-
Boundary sheet
Identity, permissions, tools, approvals, escalation paths and data boundaries, agreed before code.
-
Evaluation set and results
Real cases with known answers, the accuracy achieved, and the cases that fail.
-
The working system
On your tools or in your repository, with validation and logging from the first version.
-
Monitoring view
Runs, exceptions, accuracy and cost, visible to the owner.
-
Runbook
What to do when a step fails, a tool is down, or the numbers drift.
-
Handover and maintenance plan
A named owner, documentation, and what a month of looking after it involves.
09
Technical and risk notes
What we look at
- Integration surfaces: APIs, webhooks, exports, mailbox access, or screen-only systems.
- Identity for the automation: its own service account, its scopes, and how each tool enforces permissions outside the model.
- Structured versus unstructured inputs at each step, and the extraction schema for the unstructured ones.
- Validation rules and confidence thresholds, and what happens below them.
- Idempotency and retries: can a step run twice without a duplicate record, email or payment?
- Logging per run: inputs, outputs, tool calls, model version, who approved.
- Data flow: which providers receive what, retention, region, training use.
- The existing Zapier, Make or Power Automate flows: what they do, who owns them, where they fail.
- Cost per run, and the ceiling.
What can go wrong
- The exceptions turn out to be most of the volume, and the automated flow routes nearly everything to a person.
- A retry sends the same email or creates the same record twice.
- Extraction accuracy that looked fine on ten examples drops on the real mix of formats.
- An agent with broad access reads a document that contains instructions and acts on them. That's prompt injection: instructions hidden in content the system reads.
- The service account has admin rights because that was easiest on the day.
- Nobody is watching the exception queue, so failures are silent.
- The provider changes a model version, the classifier drifts, and nothing measured it.
- The flow depends on one person's mailbox credentials, and they leave.
10
Questions people ask
Which kinds of work suit plain automation, an AI step, or an agent?
If the input is structured and the rules can be written down, plain automation. If one step has to read something unstructured, an email, a PDF, a free-text request, add an AI step there and keep the rest deterministic. An agent is for work where the path varies each time and a person would otherwise choose it, and it needs the supervision to match. The rule: not every problem needs an agent, and most do not.
What does safely actually mean?
The system runs as its own identity with the least access it needs, and actions with consequences wait for a person until you decide otherwise in writing. Every run is logged, exceptions go to a queue someone watches, and you know which providers see which information. We put all of this in a boundary sheet before we build, so safely is a list, not a feeling.
Do you build on our existing tools or write custom software?
Both, depending on the step. Zapier, Make and Power Automate are fine for predictable steps with good connectors, and we keep them where they work. Custom code comes in for the AI step, for validation, for permissions and logging the tools cannot express, and for systems without connectors. The design note says which is which and why.
How long does it take, and how is it priced?
One workflow, from map to hand-over, typically runs six to twelve weeks, with the map and design done in the first two. Pricing is a fixed scope quoted after the map, because the map is what tells us the size. If you only want the map and the evidence, that is available on its own.
How do we know it is working?
Against numbers you chose at the map stage: minutes per case, error rate, exceptions reaching a person, cost per run. The evaluation set gives you accuracy before launch; the monitoring view gives it to you after. If the numbers drift, the runbook says what to do and who does it.
What does maintenance look like?
A few hours a month for the owner: watching the exception queue, re-running the evaluation set when a provider changes a model, and adjusting rules as the business changes. We can do that by the month, or hand it to your team with the runbook. Either way, someone is named.
Can you give examples from our kind of business?
Law and accounting firms: intake, conflict checks, document requests, engagement letters; brokerages and property managers: change requests, renewals, tenant and client correspondence; distribution and services: order changes, dispatch, invoice matching; agencies: briefs, approvals, reporting. The shapes repeat across sectors. The rules and the judgment steps are yours.
Can an AI agent use our systems safely?
Only with permissions it cannot widen, a short list of tools, approval on actions with consequences, and a log of every step. If those are in place, yes. If they are not, the honest answer is not yet, and we would start with a deterministic flow and an AI step instead.
Bring the workflow. Bring the Zapier flow nobody trusts.
We'll map it with you and say which kind of automation it needs, or whether it needs any.