Capability 02
Test vendors against your workloads, not their demo.
AI platforms, SaaS tools and cloud security vendors, run side by side on what you actually run. Scored on what it costs over three years, and on whether the vendor is still there at renewal.
Best fit: choosing between providers
The problem
what goes wrongA demo is built to succeed. Clean data, a happy path, and a sales engineer who knows every workaround. Your environment has none of those. That is why software so often underdelivers the week it meets real conditions.
There is also a survival problem. The 2026 Marketing Technology Landscape added 1,488 products over the year and retired 1,367 — about 124 in and 114 out every month. Nearly half the products removed were earning $1M–$10M. Picking a vendor is partly a bet that it still exists at renewal. Source: chiefmartec / MartechTribe ↗
What we do
four piecesRequirements that bind
Criteria written from your workloads and weighted before any vendor is contacted. A demo doesn't get to set its own test.
Hands-on bake-off
The shortlist runs the same tasks, on the same data, through the same failure cases.
Three-year TCO
Licensing, egress, integration and the staffing cost of running it. Handed over as a model you can re-run.
Viability read
Funding, customer concentration and acquisition risk. We also check how revenue is reported: some AI vendors book cloud-channel revenue gross, which flatters the headline (a worked example).
Tools & resources
how this service uses themDownloads for this service are above. All free tools run in your browser and need no email.
What good looks like
6 of 6 coveredWhat buyers are increasingly measured against here, taken from the frameworks doing the measuring. Each line names the engagement that covers it. Where nothing does yet, the row says so.
| Capability | What done properly looks like | Covered by |
|---|---|---|
| Criteria agreed before vendors are contactedProcurement practice | Scoring is written from your workloads first, so a demo cannot set the terms of the comparison. | 2.2 Bake-Off Report |
| Testing against real data and real failure casesProcurement practice | The shortlist runs your workloads, including the edge cases, not a vendor-prepared happy path. | 2.2 Bake-Off Report |
| Three-year total cost, not licence priceTCO practice | Licensing, egress, integration and the staffing cost of operating it are modelled together. | 2.3 Three-Year TCO Model |
| Vendor viability treated as a scored criterionSupply-chain risk practice | Funding stage, customer concentration and acquisition risk are scored, because a product can be retired out from under you. | 2.2 Bake-Off Report |
| A documented exit pathEU AI Act · Provider obligations | What it takes to leave — data export, integration rework, notice periods — is known before you sign. | 2.3 Three-Year TCO Model |
| Contract monitoring between renewalsVendor management practice | Pricing changes, funding events and end-of-life notices are tracked, not discovered at renewal. | 2.4 Renewal Watch |
Sample report: 808 Fit & Viability Matrix
808 demo datasetFictional demo company · sample data, not a client
Lakeshore Example Co.
A fictional 1,200-person distributor we use to show what our reports look like. Every number below is invented for the demo.
- 4 vendors on the shortlist for customer-service AI
- Vendor A picked: highest fit, mid-range three-year cost
- Vendor B kept as runner-up: strongest viability
- Vendor C dropped: viability 2 of 5
Engagements & products
status shownThe deliverables behind this capability. We publish what is live and what is still being built rather than implying a bench we don't have.
-
2.1 Entry Planned
Vendor Shortlist Teardown
Free published comparison
A public head-to-head of two or three vendors in one category, scored against stated criteria with the reasoning shown — including where the runner-up wins. Demonstrates the method rather than describing it.
-
2.2 Anchor engagement Planned
Bake-Off Report
Fixed-fee engagement · 4–6 weeks
Your shortlist run side by side against your real workloads, your data and your failure cases, scored against criteria agreed before any vendor is contacted. You get a comparison you can defend in a procurement review, and a recommendation that states its trade-offs.
Shadow Scanner The scan shows what is already in the estate, so the evaluation starts with the incumbent stack instead of a blank page.
-
2.3 Anchor engagement Planned
Three-Year TCO Model
Deliverable artifact · spreadsheet
A working cost model covering licensing, egress, integration and the staffing cost of actually running the thing. Handed over, not presented — so you can re-run it yourself when pricing changes.
Shadow Scanner Duplicate-spend findings feed the baseline, so the model starts from what you already pay.
-
2.4 Continuing Planned
Renewal Watch
Annual subscription
Ongoing monitoring of your vendor set: pricing changes, funding events, acquisitions, end-of-life notices, and a flag when a contract is coming up for renewal. Built for a market where products are retired almost as fast as they launch.
Shadow Scanner The scan defines the watch list, and re-scans keep it current as the estate shifts.
Status is kept in one place and shown as it stands. See the full catalogue across all six capabilities.
Further reading
frameworks · research · our work, on this topic- Op-edToo Fast to Fail: reading an AI vendor’s revenue before you signHow we read a fast-growing AI vendor’s numbers before a long contract.The 808 →
- GuideAdopting AI Responsibly: guidelines for procurement of AI solutionsA checklist for buying AI in the private sector.World Economic Forum ↗
- GuideGuidelines for AI procurementPublic-sector buying rules that translate well to any procurement team.UK Government (DSIT) ↗
- PolicyM-25-22: Driving efficient acquisition of AI in governmentUS federal rules on buying AI, including avoiding vendor lock-in.US Office of Management and Budget ↗
How Shadow Scanner helps
platform intelligence & risk telemetryBefore you compare new vendors, know what you already pay for. Shadow Scanner maps detected tools against sanctioned equivalents and surfaces the overlap. The evaluation then starts from the incumbent stack, not a blank page.
What you walk away with
the outcomeA scored comparison you can defend in a procurement review. A recommendation that states its trade-offs, including where the runner-up would have been the better call.
The trade-off. A real bake-off takes four to six weeks and needs your data and your people's time. A shortlist from a demo takes an afternoon. The difference shows up at renewal.
Tell us what you're running.
A person replies within one business day.