Skip to main content
AI Harness Engineering

The control layer around your AI. Guardrails your auditor can read.

Evaluations, gates, approvals and replay around the agents you already run, so what they are allowed to do is enforced before they act, not discovered after.

The market calls it harness engineering. Your compliance team will call it evidence.

The problem

Your agents run on prompts and hope.

  • The guardrails live inside a vendor platform, and you cannot show them to anyone.
  • There is no test suite for behaviour, so nobody knows what changed when it changed.
  • A change to a prompt is a change to production, with no review and no version.
  • When it goes wrong, a log is all you have, and it answers the wrong question.
  • Your risk team cannot tell the difference between an agent that is controlled and one that has not failed yet.

Reliability is an engineering problem. Permission is an evidence problem. This is both.

How we work

In. Build. Leave. Prove.

One job, four weeks, in production. Then it keeps proving itself.

  1. In

    The behaviours that matter and the rules they have to respect, written down first.

  2. Build

    The control layer around one agent, then the rest, wired into your CI.

  3. Leave

    Your engineers run it in CI and own the rules and the evaluations.

  4. Prove

    Every decision passed the gate or was blocked, and the record says which.

What we build

What you are left with.

  • An evaluation suite for the behaviours that matter, run like tests.
  • Checks against your rules that run before an action takes effect.
  • Approval gates with named people at the steps that carry consequence.
  • Prompt and policy versioning, with a sign-off recorded against each change.
  • Replay of any decision, to the same answer, any day.
  • All of it wired into the CI your engineers already use.
Built for the regulator's questions

Can you prove this agent has only ever done what you allowed?

  1. What the AI is allowed to do is written down first.

  2. Every decision is checked against those rules before it happens.

  3. A named person signs it off. The check produces the evidence; a person judges.

  4. You can replay the whole history any day and get the same answer.

  5. If anyone changes it later, it shows. Patent pending, UK application GB2620101.2.

That is what a deployment leaves running for your job. In the accessibility product today, a person on your team accepts every finding before it reaches your record.

Who this is for

Who this is for.

  • Engineering and risk leaders whose agents are already in production, or about to be.
  • Teams who have to answer a model risk or operational resilience question about AI.
  • Anyone who has been asked to show the guardrails, and found they live in someone else’s platform.

Not for

Teams without an agent yet. Start with Custom AI Agents, and this comes with it.

Why us

Two people, on every call and in your standup.

Simon Milner, Founding Architect

He designed the record: what the AI is allowed to do, checked before it acts, and replayable afterwards. Twenty-five years in Silicon Valley before that.

Jason Crispin, Founder

He owns the customer side of every deployment: what the job is, what it is worth, and that it lands. He is on the first call and every one after.

Patent pending, UK application GB2620101.2. Meet the team

How it runs

Four weeks, then it keeps proving itself.

  1. Week 1

    The baseline

    What the job is, what allowed means for it, and who signs. Written down before anything runs.

  2. Weeks 2 to 4

    The build

    Our engineer works in your codebase next to your developers. The old way and the new way run side by side.

  3. Week 4 on

    The proof

    Every decision checked and recorded. Replay it any day. We maintain it, or you run it without us.

What you keep

  • The control layer, running in your CI.
  • The evaluations and the rules, owned by your engineers.
  • The code, assigned to you in writing.
  • The record of every decision, replayable any day.
Questions

What people ask.

Is this AI governance?

Is this AI governance?

It is the engineering that makes governance provable. Your policy says what the agent may do; this enforces it at the moment the agent acts, and records the check. A policy document cannot do that.

Does it slow the agent down?

Does it slow the agent down?

The check runs before the action, so it costs what a check costs. What it removes is the slower path: an incident found later, and the week spent reconstructing what happened.

Does it work with our model provider?

Does it work with our model provider?

Yes. The checks sit around the agent rather than inside a model, so you can change provider or model without rewriting the rules.

Who owns the rules?

Who owns the rules?

You do. They are written in your words, versioned, and changed by your team with a sign-off recorded against each change.

We already have logging. Why is this different?

We already have logging. Why is this different?

A log tells you what happened. It cannot tell you whether each action was allowed, or whether the log itself has been changed. The check happens before the action, and the record can be replayed to the same answer.

What does it cost?

What does it cost?

Scoped on the call, because it depends on how many agents you run and what they touch.

Which agent would you put under this first?

Thirty minutes with Simon and Jason. Bring the agent and the rule it has to respect, and you leave knowing what it would take.