September 2026

Reference · vendor-agnostic

The enterprise AI stack.

Six layers, what to look for in each, and what usually goes wrong. One test runs through every layer: can you see what it is doing, and can you take it out cleanly? Most AI programmes get assembled fast, one layer at a time. So the failures show up at the seams, not inside any one product.

The layers

top to bottom

A decision at one layer constrains the layers around it. A model choice locks in an orchestration pattern. A data platform decides what an agent can safely see. Read them together.

01

Applications

The software people touch, where an AI output becomes a decision, a document or an action. Less often a chat window now. More often an existing business application that has quietly grown agentic behaviour.

What to look for

Whether the vendor lets agents drive the system directly rather than only through its own interface, and whether you can see what an agent did on your behalf.

What usually goes wrong

A capable model wired into an interface nobody uses. Or an agent allowed to act in a system that keeps no record of what it changed.

We cover this in Implementation & integration · AI strategy & roadmapping

02

Delivery & lifecycle tooling

The toolchain used to build, test, version, evaluate and monitor AI capability — the equivalent of CI/CD for systems whose output is probabilistic rather than deterministic.

What to look for

Evaluation you can run repeatedly against your own cases, versioning that covers prompts and tools as well as code, and monitoring that reports cost and failure alongside accuracy.

What usually goes wrong

Prototypes that cannot be reproduced, and changes shipped on the strength of a few good demos because there is nothing to test against.

We cover this in Implementation & integration · Team enablement & training

03

Models

The reasoning layer. Frontier APIs, mid-tier models, open-weight models you host, and small task-specific ones — usually several at once rather than a single choice.

What to look for

The smallest model that meets the requirement, and whether your data trains anyone else's system. If a system downstream acts on the output, the model must return structured results against a schema, not prose someone parses. Plan to change model within eighteen months.

What usually goes wrong

Standardising on one provider, then discovering the integration is written against its conventions and switching means a rewrite rather than a config change.

We cover this in Vendor & tool evaluation

04

Orchestration & runtime

Where agents plan, call tools, hold state and act across systems — increasingly several agents handing work to each other. The layer that turns a model from something that answers into something that does. Demos look alike here; operations do not.

What to look for

Purpose limits per agent, enforced spend and rate caps, human approval at irreversible steps, a kill path someone has tested, and action logs that survive a review. When agents delegate, each hand-off must keep the original authority attached.

What usually goes wrong

Controls added after an incident. This is where an unbounded loop becomes a bill and an agent inherits whatever credentials it was handed. With several agents, nobody can say which one acted, or on whose request.

We cover this in Agentic AI engineering

05

Data platform

The governed supply of structured and unstructured material an agent reasons over, plus the access paths and permissions that decide what it may see.

What to look for

Real-time access rather than nightly exports, and permissions that follow the user through to retrieval. Shared definitions for the business terms agents use, so "customer" or "active account" means one thing everywhere. A record of what was retrieved to produce each answer.

What usually goes wrong

An agent that can read more than the person asking it — the most common way sensitive material leaks without any breach. And two agents giving different answers because two systems define the same term differently.

We cover this in AI strategy & roadmapping · Governance, risk & compliance

06

Infrastructure

Compute, storage and network underneath it all. For most buyers this is now a question of inference cost and latency, not training capacity.

What to look for

Cost per inference at your real volumes, latency under concurrency rather than in a benchmark, and where the workload may run for residency reasons. Cost and reliability on one dashboard, watched by one team, since a retry storm is both an outage and an invoice.

What usually goes wrong

Pricing modelled on pilot volumes. Inference cost scales with usage, so the bill arrives precisely when the thing starts working.

We cover this in Vendor & tool evaluation · AI strategy & roadmapping

Evaluating across the stack

8 criteria

The same criteria apply at every layer. That is what makes a comparison across products mean something. We agree and weight them with you before any vendor is contacted. A demo shouldn't set the terms of its own test.

  1. 01

    Functional fit to real use cases

    Does it do the specific job you need, on your data, at your volumes — not the job it demos well.

  2. 02

    Operability in production

    Can you deploy, monitor, debug, recover and upgrade it with the team you actually have.

  3. 03

    Deployment & sovereignty flexibility

    Does it fit your hosting model and data residency obligations without architectural compromises.

  4. 04

    Governance, security & compliance controls

    Does the product enforce permissions, budget caps and audit trails itself, or expect you to build them.

  5. 05

    Integration & ecosystem fit

    How well it connects to what you already run, and what it costs to disconnect it later.

  6. 06

    Total cost to run

    Licensing plus inference, egress, integration and the staffing cost of operating it for three years.

  7. 07

    Vendor viability

    Funding, customer concentration and acquisition risk — because a retired product is an unplanned migration.

  8. 08

    Safe to remove

    Can you switch it off, undo what it did and take your data with you, without the vendor's help. Ask before you sign, not at renewal.

How we run an evaluation Score your own readiness first

Start from what you actually run

inventory first

Every organisation we work with can describe its intended stack. Far fewer can describe the one that exists: the consumer AI logins, the personal-tier developer tools, the unmanaged GPU box somebody spun up. Shadow Scanner inventories what is running across these layers and who uses it. That is the honest place to begin.

Several research firms and platform vendors publish layered models of the AI stack. Read them alongside this one. The definitions, failure modes and criteria above are ours, written from engagements rather than summarised from a report.

Which layer do you fix first?

Tell us what you're running. A person replies within one business day.

Talk to us