September 2026

Capability 05

Give AI the ability to act, with limits that hold.

Agent demos converge. Operations don't. We design and harden agent workflows with spend caps, approval gates and full action logs built in from the start — and a kill path someone has actually tested.

Best fit: before an agent can act

One-page service sheet preview

Downloads

take it to the meeting

Service sheets follow one standard across all six capabilities. Engagement status on each sheet is taken from the live catalogue.

The problem

what goes wrong

An agent that answers questions is a low-stakes system. An agent that can spend money, write to a database, send mail or call an API is a different category of risk. Most organisations move from the first to the second without changing their controls.

By then the models are close enough that capability rarely decides anything. Operability does. Can you limit what an agent may do? Can you stop one mid-run, and prove what it did? If not, the blast radius is whatever credentials it was handed.

What we do

four pieces
01

Scoped agent design

What each agent may touch, and what it must never touch, defined before anything is built.

02

Hard limits

Spend caps and rate limits enforced at the platform, so a loop can't run up a bill overnight. A kill path that has been tested, not assumed.

03

Human-in-the-loop gates

A person approves at the points where a wrong call is expensive or can't be undone.

04

Full action logging

Every action traceable to an agent, a trigger and an authority. An incident review has something to review.

What good looks like

7 of 7 covered

What buyers are increasingly measured against here, taken from the frameworks doing the measuring. Each line names the engagement that covers it. Where nothing does yet, the row says so.

CapabilityWhat done properly looks likeCovered by
A purpose limit per agentNIST AI RMF · Govern What each agent may touch, and what it must never touch, is defined before it is built. 5.2 Guardrail Reference Build
Spend and rate limits that bindOWASP LLM · Unbounded consumption Caps are enforced at the platform, so a loop cannot run up a bill overnight. 5.2 Guardrail Reference Build
Human approval at irreversible stepsNIST AI RMF · Manage Anything expensive or hard to undo pauses for a person, by design rather than by convention. 5.2 Guardrail Reference Build
A kill path that has actually been testedIncident response practice Someone has stopped an agent mid-run in a drill, and knows how long it took. 5.3 Agent Incident Drill
Full action attribution after the factOWASP LLM · Excessive agency Every action is traceable to an agent, a trigger and an authority, so a review has something to review. 5.2 Guardrail Reference Build
Hand-offs that keep their authorityOWASP LLM · Excessive agency When one agent delegates to another, the original request and its limits travel with the work, so the log still shows who asked. 5.2 Guardrail Reference Build
Input handling that assumes hostilityOWASP LLM · Prompt injection Content an agent reads is treated as data, never as instructions it may follow. 5.1 Agent Risk Assessment

Sample report: 808 Blast-Radius Chart

808 demo dataset
DATASYSTEMSTOOLSAGENTNEEDEDREACH, NOT NEEDEDCUT FIRST →red marks
Lakeshore Example Co. · Fictional demo company · sample data, not a client

Fictional demo company · sample data, not a client

Lakeshore Example Co.

A fictional 1,200-person distributor we use to show what our reports look like. Every number below is invented for the demo.

  • 1 invoice-matching agent mapped
  • 7 reach points across tools, systems and data
  • 3 reach points it has but does not need: cut first
  • Kill path drill: 11 minutes to stop mid-run (target: under 2)

Engagements & products

status shown

The deliverables behind this capability. We publish what is live and what is still being built rather than implying a bench we don't have.

  • 5.1 Anchor engagement Planned

    Agent Risk Assessment

    Fixed-fee engagement · 2–3 weeks

    Where agents are already acting outside managed boundaries, what credentials they hold, and what the blast radius looks like today. Most organisations are further along than their policy assumes.

    Shadow Scanner The strongest fit of any engagement. Unmanaged model endpoints, personal-tier developer tools and assistants joining calls off-network are all direct scan output.

  • 5.2 Anchor engagement Planned

    Guardrail Reference Build

    Productised build

    A reference implementation you keep. Scoped permissions, spend caps and rate limits. Human approval at the expensive or irreversible steps. Full action logging, and a kill path that has actually been tested.

    Shadow Scanner Scan findings set the initial scope — what each agent may touch, and what it must never touch.

  • 5.3 Continuing Planned

    Agent Incident Drill

    Workshop · half day

    A rehearsed exercise: an agent misbehaves, and your team has to detect it, stop it, and reconstruct what it did. Most teams have never tested the kill path. Better to find that out here.

    Shadow Scanner The drill runs against your real agent inventory, so the scenario is yours rather than a generic one.

Status is kept in one place and shown as it stands. See the full catalogue across all six capabilities.

How Shadow Scanner helps

platform intelligence & risk telemetry

Most environments already run agents — just not the ones anyone approved. Shadow Scanner finds where AI is acting outside managed boundaries. That is the honest starting point for building agents that stay inside them.

More about Shadow Scanner

What you walk away with

the outcome

Agents that do useful work inside boundaries you set, with a record of what they did and a way to stop them that you have rehearsed.

The trade-off. Approval gates slow the workflow at exactly the steps that matter. We put them only where a mistake is expensive or permanent, and say which steps run unattended.

The five stages

Tell us what you're running.

A person replies within one business day.

Talk to us