Skip to content

How we deliver

Your workflow running under your rules, with proof it beats what you do today.

Getting there takes four stages. You can stop after any of them and keep what you have, and the first one tells you whether the rest is worth doing — counted from your own history rather than promised from a demo.

Five multiple-choice questions. A person replies within one business hour.

Before anything is built

In month three someone will ask whether it worked. Today you could not answer.

Ask a room why last year’s pilot died and you will hear about the model. Ask what the work cost before the pilot started — how long people waited, how many requests never got an answer, how often the same question came back — and the room goes quiet. Without those numbers the go or no-go is settled by whoever is most senior, not by what happened.

A fortnight closes that hole. The numbers come out of your own history before a line of anything is built, and you keep them either way. If the automatable share turns out to be too small to bother with, that is a good outcome for a fortnight and you will have it in writing.

There is a second reason, less flattering to us. A baseline is the only honest way for you to hold us to account. It is very easy to make an agent look good in a demo and very hard to make it look good against a number that was recorded before we arrived.

5%
of custom enterprise AI tools reach production with measurable impact
MIT Project NANDA, preliminary study, 2025
85%
say gaps in traceability or explainability have delayed or stopped an AI project
Dataiku-commissioned Harris Poll, 600 CIOs
25%
have full real-time visibility into the AI agents they already run in production
Same survey

Those are not here to frighten anybody. They are why the first thing you get from us is a set of numbers about your own operation rather than a workspace and a login.

The four stages

Stop after any stage and you are still better off than you were.

Nothing here bills by the hour and nothing obliges you to the next stage. Each card says what is true for you at the end of it, what staying as you are costs in the meantime, and what it takes from your side.

Diagnose · The X-RayPreview

You find out what is genuinely automatable here, counted from your own history.

Two to three weeks

We read your own conversations and operations, privacy-first, and hand you the numbers: how long people wait, what never gets answered, what repeats, and exactly which share an agent could take — and which actions would still need a person.

What staying as you are costs

Another quarter of deciding by anecdote. Then in month three someone asks whether the AI helped and by how much, and the honest answer is that nobody counted before it started.

What happens

  1. Days 1–2Nothing touches production. You send a read-only export of the workflow's own history; no write access to anything, and no change to how the work runs while this happens.
  2. Days 3–7The counting. How long people wait for a reply, what share of requests never got an answer at all, which questions come back week after week, and when the queue really peaks rather than when you think it does.
  3. Days 8–11The sorting. Every recurring request becomes the list of actions it implies, and every action lands on the risk tier table: read, draft, reversible change, external, irreversible. This is where you see which part an agent could take and which part has to stay with a person.
  4. Days 12–14The read-out. Ninety minutes with your team, then the written baseline. Yours to keep whether or not we ever work together again.

What you have at the end

  • Baseline: response times, unanswered share, repeat questions
  • The automatable share, counted from your own history — not estimated
  • An approved-answer backlog you own, whether or not you continue
  • The action list split by risk tier: what an agent may do, what a person must

What it takes from your side

  • Six to twelve months of the workflow's own history, exported read-only
  • One person who knows how the work is really done, for about three hours in total
  • Ninety minutes for the read-out, with whoever makes the decision in the room
This has run end to end on one customer's own history. The tooling is still partly bespoke and is being generalised, so the scope is agreed with you before anything starts.
Build, fast · Governed Agent BootcampDesign partner

One real workflow is live in a week, with the rules agreed in writing before it runs.

One week, on site or with your team

One real workflow live in your own workspace, with its approval policy, its audit trail and a run-book — then a go or no-go you can defend to whoever asks.

What staying as you are costs

The arguments about what an agent may do happen either way. Put off, they happen after go-live, with a customer already affected and no written policy anyone can point at.

What happens

  1. Day 0The workflow is chosen and written down with its actions listed, before the week starts. If that list cannot be agreed, the week does not start.
  2. Day 1Tiers. Every action in the workflow goes onto the tier table with your team in the room. The arguments happen here, which is the point: they cost a morning now or a quarter later.
  3. Day 2The workspace. The agent, the tools its policy lets it see, the rule for anything above the line, and the person who receives those approvals, by name.
  4. Day 3Real data. The workflow runs against your own records rather than a sample, in a workspace that belongs to you.
  5. Day 4Break it, in front of you. Including the two tricks that work on other systems: telling the agent the customer has already approved, and telling it to raise its own limit.
  6. Day 5The run-book, and an honest go or no-go you can hand to whoever asks for it.

What you have at the end

  • One workflow chosen on day zero, in writing
  • Its risk tiers and approval rules, configured with your team
  • The audit trail, running
  • A run-book, and an honest go / no-go

What it takes from your side

  • Four half-days from the person who actually owns the workflow
  • Read access to the system of record, and a test account for anything that writes
  • A named approver — a person, not a shared inbox
Every piece of this week is work we have done. The five-day packaging is new, so the first cohorts run with the people who built the product at the table rather than a delivery team reading a script.
Build, full · Governed PilotPreview

A complete system proves itself against your real traffic before a customer ever sees it.

Six to eight weeks

One complete system live on the operating system: agents, approval rules, your own branded console, a shadow run against your real traffic, and a scorecard you can put in front of a board.

What staying as you are costs

Go straight to live and your customers are the test set. The first thing you learn about a bad answer is that somebody already received it.

What happens

  1. Week 1Scope, tiers and the approval policy, agreed in writing — including the part everyone skips, which is what result would make this a no.
  2. Weeks 2–3Build. Agents, rules, interfaces and connectors, with your brand and your domain on the front and the governance behind it.
  3. Week 4Shadow. The agent answers every real item as it arrives, and not one of its answers reaches a customer. You read them.
  4. Week 5The scorecard. Measured against your own history and against a held-out slice the agent has never seen, so it cannot score well by having memorised the answers.
  5. Week 6Go or no-go, with the scorecard on the table. If it is go, traffic moves in slices rather than all at once, and a person watches the first slice.

What you have at the end

  • A system, not a demo: agents, rules, interfaces, connectors
  • Your brand on the front, our governance behind it
  • Shadow run measured against your real traffic before anything goes live
  • A scorecard, and the decision that follows from it

What it takes from your side

  • A named approver with real authority, not a delegate who has to ask
  • A data path to the system of record, agreed with whoever owns security
  • Someone senior who is willing to read a scorecard that says no
One system runs this way today, for one customer. On that deployment our own gate said not ready and it did not go live on the date we had planned. Shadow mode and the scorecard are delivery practice we run for you; they are not yet switches you flip yourself.
Operate · Managed operationsPreview

The agents keep working when the business changes, and a named person answers the phone.

Ongoing, reviewed quarterly

We run the agents: monitoring, the learning loop, policy tuning, cost and model routing, and a monthly governance report with a named operator who answers the phone.

What staying as you are costs

An agent nobody tunes drifts. The questions it could not answer pile up unanswered, the policy ages against a business that moved on, and the first sign of either is a complaint.

What happens

  1. Every dayMonitoring, with a person on the other end of it. When something fails it falls back to a human, never to silence.
  2. Every weekThe learning loop. Questions nobody could answer become proposed answers, and a person on your side approves or rejects each one. Nothing joins the approved set without that yes.
  3. Every monthA governance report: what ran, what waited for a person, what was refused, what it cost, and what we propose changing.
  4. Every quarterA policy review. Tiers move deliberately or not at all, and loosening a rule is itself a high-tier action with its own approval.

What you have at the end

  • Monitoring and incident response
  • The learning loop: unanswered questions become proposals you approve
  • Monthly governance report
  • A named operator, not a ticket queue

What it takes from your side

  • One owner on your side who can approve proposals
  • Thirty minutes a month for the report
  • A call when the business changes, because the policy has to change with it
Operations run this way for one customer and for our own company. The monthly governance report is assembled by a person from the ledger today; the self-serve version of it is not built.

Whichever stage you stop at, you leave with the same first artefact: your own actions, sorted by what an agent may do without asking.

  • T0Read
  • T1Draft
  • T2Reversible change
  • T3External — a named person approves
  • T4Irreversible — a person only
  • T5Change the rules — owners only

The full table, and how each tier is enforced, is on the governed autonomy page.

What you can hold us to

Two promises you can rely on today, and two you should not yet.

A promise our own ledger or our own billing cannot enforce is worse than no promise, and it is worse for you, because you are the one who would be relying on it. So here are all four, with the two that are real marked as such and the two that are not marked plainly.

  • You see the scorecard before your customers see the agent

    Yours today

    Your agent runs in shadow against your real traffic, and you read what it would have said, before it speaks to anyone. If the week or the pilot does not end with your workflow running under your own approval policy, you do not pay for it.

    Yours today. It is delivery practice we have run, and it is the reason our first deployment did not go live on its planned date.

  • Leave inside sixty days with everything, including the record

    Yours today

    Cancel in the first sixty days and you walk out with every workflow, every policy and every ledger record from your workspace. Nothing you paid for is held hostage in our software.

    Yours today. That export is produced as an operator task rather than a button you press; the button is on the roadmap and the commitment is not waiting for it.

  • A credit if a consequential action ever runs unrecorded

    Not yet

    Anything at tier 2 or above is written down today with its tier, its check and its approver, so you can already show an auditor what ran and what was stopped. The commitment still to come is money back when a record is missing.

    The recording half is real and running today, and it is the half you rely on. The credit half is not a promise yet, because there is nothing to credit until sign-up and billing exist. It becomes written when it becomes enforceable.

  • You are never billed above a ceiling you set

    Not yet

    The intended commitment is simple: set the number, and an overshoot caused by our metering is ours to absorb.

    Not a promise yet, and two things ship before it is: a workspace spending ceiling is in preview and not yet enforced before an agent runs, and there is no metering to cap. What does work today is per-run limits, loop detection and provider circuit breakers, and every trip is written to the ledger.

The split is the reassuring part, once you look at it. What you actually depend on — the shadow run, the exit with your data, the record of every consequential action — is running today. What is missing is the money half of two promises, which needs billing that does not exist yet. You get to check that here rather than find it at a renewal.

The rules that do not move

Six things that stay true on the week you are in a hurry.

  • Your approved answers stay word for word yours. We build the system, not your content.
  • Your numbers come from real traffic before go-live, not from a sample somebody chose to flatter the result.
  • Shadow before live. Always, and especially when you are in a hurry.
  • People approve money. Every time, by name.
  • When something fails it falls back to a person, never to silence.
  • The decisions that are yours — prices, policies, payment instructions — are surfaced to you, not written for you.

When we say no

Four situations where you should not hire us.

Not modesty. Each one has a reason, and each one saves you a quarter you were otherwise going to spend finding it out.

  • We decline

    You want a chatbot, and you do not want the measurement.

    Without a baseline there is nothing to hold the result against, and the governance around it becomes ceremony you are paying for rather than a limit doing work. If what you need is a bot on the website for a monthly fee, that market is well served and we will name a product on the call.

  • We decline

    Your procurement gate is a certificate, this quarter.

    We hold no third-party attestation today, and no architecture argument gets you past a gate written as a certificate with a date on it. You would spend three weeks and lose it in security review. Come back when the gate is a security review, which we answer in detail and in public, or bring us in behind a vendor who already holds the certificate.

  • We decline

    You want an agent to take an action a regulator reserves for a licensed person.

    Where the rules require a licensed human to make the recommendation, sign the filing or give the advice, we will not build an agent that does it — not with a disclaimer, and not with a human in the loop who is really clicking through. Everything around it, yes: the intake, the preparation, the evidence trail, and the handoff to the person licensed to decide.

  • We decline

    Nobody will put their name on the approvals.

    The system holds irreversible work for a named person. If no one will be that person, the work simply queues, and in eight weeks the conclusion in your business will be that agents do not work here. That argument is cheaper to have now than to lose later to a queue.

If one of those is you, say so in the first email. You will hear it on the call rather than after a scoping exercise, and where we know a better answer we will name it.

Straight answers

What people ask before they sign.

Because the decision at the end is not a judgement call. The agent runs against your real traffic without answering anyone, and it is scored against your own history and against a held-out slice it has never seen, so it cannot look good by having memorised the answers. If the score does not clear the bar you set in week one, the answer is no, and you will have spent weeks rather than quarters finding that out.

Find out what is automatable here before you build anything.

A fortnight, and you end up holding a baseline of your own operation — including the version of it where the answer is that you do not need us.

Five multiple-choice questions. A person replies within one business hour.