We assess one agent workflow, then hand your engineers the control files that bind it. It runs on questions your own staff answer. We never touch your systems.
One domain below threshold decides the verdict, whatever the weighted total says. Figures illustrate the instrument, not a customer result.
Ask what a given agent is permitted to do, and up to what limit. Usually the answer lives in someone's head.
One verdict, traced end to end. The agent reaches an MCP server, which authenticates as a non-human identity, which carries a scope, which permits a refund, which needs a control, which produces the evidence an auditor asks for. Grey edges are relationships the assessment records but this verdict does not turn on.
Which actions, under whose identity, up to what limit. In most estates that has never been written down, and the answer sits with whoever built it.
An auditor can only test a boundary that was declared somewhere. Your second line has the same problem.
The agent does something nobody decided it could do. The model behaved. The permissions did not exist.
We take one workflow at a time. Five things come back, and all five are yours to keep.
It arrives as one folder with five numbered children, in the order the work happens. The report says where you stand. The tracker turns that into something a programme manager can run. The summary is what a board reads before it funds anything. The packs are what your engineers apply. The evidence register is what makes all of it defensible a year later. Four different people open four different files, and none of them has to call us first.
The score, a scorecard per domain, every critical and high finding, where the workflow is exposed, and the order to fix it in. Run it again next quarter and the number is comparable.
The same findings as a tracker, with an owner, a priority, a target window and the dependencies against each one. It turns the report into a nought-to-ninety-day programme somebody is accountable for.
One page. Where the posture stands, which risks matter, and the decisions currently waiting on someone. Enough for a board to fund the work or decline it.
The core agent security pack, plus only the use-case packs your assessed agent's authority reaches. Every pack carries a manifest, a README, policy in YAML, validation tests and its evidence requirements.
An evidence register, and a folder for test output and review records. This is the part that makes the controls provable a year later, when somebody asks.
An agent that only reads a knowledge base and an agent that can issue a refund do not need the same controls, and pretending they do is how a control set becomes shelfware — too heavy for the harmless workflow, and resented by the team that has to implement it. So the bundle is layered instead. One core pack is the baseline every assessed agent receives. Above it sit overlay packs, and each is triggered by a capability the agent actually holds: writing to a case, messaging a customer, moving money, reaching a privileged system. If your agent cannot do the thing, you do not receive the pack that governs it.
So the bundle is different for each workflow you assess, and the difference is not a pricing tier. It is a statement about blast radius — the same statement your risk function is trying to make when it asks how far a given agent can reach.
| Control pack | Covers | Case-triage agentread-only | Refund agentmoves money |
|---|---|---|---|
| core-agent-security-v1 | The universal baseline. Every agent, every domain. | ships | ships |
| support-read-only-v1 | Search, retrieval, knowledge, case summary | ships | ships |
| support-case-write-v1 | Ticket notes, fields, approved state changes | not shipped | ships |
| customer-communications-v1 | Email, chat, SMS, payment-status messages | not shipped | not shipped |
| payments-support-refunds-v1 | Refunds, credits, chargebacks, adjustments | not shipped | ships |
| privileged-operations-v1 | Treasury, settlement and admin actions | not shipped | not shipped |
| Packs shipped | 2 of 6 | 4 of 6 |
The core pack ships every time. The overlays switch on only where the agent's authority actually reaches, so a read-only triage agent receives two packs and a refund agent receives four. You do not get sent the packs you have no use for.
No portal to log into and no licence that lapses. A directory, handed over once:
Open any pack and the shape is the same: a manifest, a README, the policy itself as YAML, the validation tests that prove the policy holds, and the list of evidence you are expected to retain. That last file is the one an auditor asks for, and the one most control libraries do not ship.
The packs specify the controls. They are not live until your engineers configure them, test them, and keep the evidence. We say so in the deliverable itself, because a pack sitting in a repository protects nobody and we would rather you heard that from us than from an auditor.
Nothing to integrate and nothing to install. It runs on answers your own people give.
Not during the assessment, not afterwards. We built it this way so your risk function can approve the engagement on its normal path, without raising an exception for third-party access.
Every control is answered from fixed options rather than free text — the answer itself, the enforcement type, which function owns the control, and whether evidence exists — with room for a 250-character description of the current state and nothing beyond that. The form is deliberately narrow. A narrow form is one your own staff can complete honestly in an afternoon, and one your legal team can read through in ten minutes.
Because it scores control maturity, it never needs the operational detail. No payment-card or bank-account data. No tokens, CVV, credentials, API keys or certificates. No exact monetary thresholds or fraud-rule parameters. No full system prompts, no production configurations, no unredacted logs or screenshots, no internal-only URLs, and nothing your own classification scheme calls Confidential or Restricted. Where a control turns on evidence, that evidence is referenced by name, or reviewed live with your people in the room. It does not come to us, and the form says so on every question.
Two things cross the line, in opposite directions. Everything on the bottom row stays on your side of it, for the whole engagement.
One agent workflow: what it is allowed to do, and what stops it. Whoever runs that workflow can answer without preparing.
The same instrument every time, so this quarter's score can be set against last quarter's. Where you cannot show evidence, the score records it as unevidenced. We do not give credit for intent.
Report, remediation plan, executive summary, the control packs your agent's authority requires, and the evidence register. Deployment is your engineers' work, through your own change process.
Scopes widen. Someone adds a tool. The re-score says what moved since the last one.
Every control we specify carries the clause it answers to. Your compliance team can check the mapping before anyone signs anything.
One workflow, assessed and handed back as a configuration pack. The first one costs nothing. If it turns out your estate is in good order, we will say so and leave you alone.
We will not ask for credentials or system access at any stage, and the first assessment costs nothing.