Four manifestos, one lineage — Agile (2001), Digital (2016), Confidence (2026), under Assurance.
Umbrella · AI-Assurance (Governance) v1.0

The Manifesto for AI Assurance

Agile changed how we build. Digital changed what we build and how we govern it. Confidence Engineering changed how we know it works. AI Assurance holds all three to one accountable thread — from intent, through delivery, to evidence of behaviour in production, to a named human or agent who is accountable.

6 values · 12 principles · one operating rule · sits over Agile (2001), Digital (2016) and Confidence (2026)

The shipping question

Every era of software has had one question at its core. Agile asked are we building the right thing, fast enough to find out? Digital asked is the business adapting as fast as the technology? Confidence Engineering asks do we know enough about this system to ship it? AI Assurance asks the question those three were always converging on:

Do we know enough about this system to ship it — and who is accountable for that answer?

AI Confidence Engineering is the discipline that answers the shipping question with evidence. AI Assurance is the operating model that produces that evidence. Governance is the accountability that consumes it — and turns it into a decision someone signs.

Why now: the asymmetry

Generative AI is driving the cost of producing software toward zero. The cost of verifying it is not falling with it. The Confidence Asymmetry Principle holds that generation and verification obey different and diverging cost laws: generation cost falls with model capability and tooling, while verification cost is bounded below by floors no model can repeal — undecidability, execution, and interaction.

Worse, when AI generates both the code and the tests of that code, their blind spots correlate. A passing verdict from a verifier that shares the generator's blindness carries almost no information. Evidence must come from outside the model.

The consequence is an inversion: the scarce resource of the AI era is not generation — it is justified confidence. In an economy where code is nearly free to write, almost all of the cost will be spent finding out whether the code is any good, and almost all of the risk will sit with whoever signs the answer.

Three disciplines, one accountable thread

AI Assurance Blueprint

the operating model

What must be tested?

Fifteen axioms of testable trust — accuracy, consistency, safety, robustness, auditability, appropriate uncertainty, bias & fairness and their peers. System types (generative, RAG, predictive, recommender, agentic, MCP), each with its own failure profile. Two pillars: validate the AI logic and outputs; validate the end-to-end experience. The 0–2 behavioural rubric. A lifecycle: Scope → Design → Run → Score → Report → Monitor. Regulatory surfaces: EU AI Act, ISO/IEC 42001, NIST AI RMF.

AI Confidence Engineering

the epistemics

How do we know the evidence justifies trust?

The Complexity Gap diagnostic. Five evidence layers: Intent → Verification → Exploration → NFR → Runtime. The recursive standard: every metric is itself held to a confidence standard. The Two Unknowns: independent evidence when AI tests AI. The Confidence Engineer: a named owner of the shipping question.

AI-Assurance Governance

the accountability

Who is accountable, at what autonomy, under which controls?

Human-in-control governance. Autonomy tiers that scale oversight to delegated authority. RACI over every agentic action. An audit trail from intent to outcome. Assurance evidence is what governance consumes — no sign-off without evidence, no evidence without an owner.

None of the three is sufficient alone. An operating model without an epistemic standard becomes activity — dashboards green, and still no honest answer to the shipping question. An epistemic standard without an operating model is philosophy with no surface. And both together, without governance, produce evidence that nobody signs for.

We value

Through assuring real AI-infused systems — probabilistic, continuously generated, increasingly autonomous — we have come to value:

Continuous assurance over point-in-time approval

A sign-off describes the system as it was at the moment of signing. The system that reaches your users is a different system: models update silently, corpora drift, agents recompose, attackers adapt. Assurance is a standing property re-earned with every change — not a gate that was once passed. The audit that matters is the one still running.

Evidence chains over isolated artefacts

A test report that cites no intent, a metric that feeds no decision, a review that connects to no release — these are artefacts, not assurance. Evidence earns its name by being chained: from declared intent, through verification, exploration and operational signal, to the decision it changed and the person who made it. A green dashboard is a claim. A chain of evidence is an argument.

Accountable ownership over delegated inspection

Assurance cannot be outsourced — not to a QA department, not to an auditor, not to the model that wrote the code. Whoever changes the system answers for what is known about it. In agentic systems this becomes architectural: every tier of machine autonomy is matched by a named human accountability, and no agent signs for itself.

Outcome integrity over output compliance

A system can meet every specification and still fail the person using it. Compliance says the outputs matched the spec; integrity says the outcomes held up in the world — for real users, under real conditions, including the conditions nobody specified.

Transparent uncertainty over manufactured certainty

Probabilistic systems do not offer certainty, and honest assurance does not manufacture it. We report distributions, not point estimates; intervals, not verdicts; what we do not know alongside what we do. Hidden gaps become incidents. Declared gaps become decisions.

Assurance as design over assurance as inspection

You cannot inspect confidence into a system that was not built to yield evidence. Observability, traceability, evaluability and reversibility are architectural properties, decided when the system is designed — not qualities a test phase can retrofit. The cheapest evidence is the evidence the system was built to produce.

That is, while there is value in the items on the right, we value the items on the left more — for systems that are generated faster than they can be read, and that act with more autonomy than they can explain.

The twelve principles

Intent

  1. 01

    Assurance begins at intent. You cannot prove a system does the right thing downstream of never defining what right means. Intent — declared, versioned, testable — is the first evidence artefact, and every other artefact chains back to it.

  2. 02

    Every claim carries its evidence, its freshness, and its confidence. “It works” is not a claim; it is a slogan. A claim states what was observed, when, under what conditions, and how much trust the observation deserves — a confidence interval, not a checkbox.

  3. 03

    Prefer fewer, better questions over more, cheaper checks. Ten thousand passing tests answer one question ten thousand times. Assurance effort goes where a different answer would change the shipping decision.

Evidence

  1. 04

    Assurance flows continuously — it is not a phase, a gate, or a department. It runs where delivery runs: in the pipeline, in production, at the same cadence as change.

  2. 05

    Risk directs effort; coverage merely records it. Allocate verification to where failure is expensive, irreversible, or silent — not to where checks are easy to write. In AI systems the risk map itself decays, so risk identification is continuous too.

  3. 06

    Production is the primary source of truth; everything before it is rehearsal. Synthetic evaluation is how we prepare; operational reality is how we know. Monitoring is not an afterthought to testing — it is the senior partner.

  4. 07

    Automate the evidence, never the judgement. Machines gather, score and trend at scales no human can; humans own what the evidence means and whether it suffices. Evidence about a generator must be independent of the generator — a verifier that shares the author's blind spots produces reassurance, not evidence.

System

  1. 08

    Governance is a product: usable, versioned, and tested. Controls that obstruct get bypassed; policies nobody can execute are decoration. Ship governance the way you ship software — with users, feedback, and releases. Unusable governance is a defect.

  2. 09

    Assure whole systems. People, processes, models, prompts, tools and services compose into behaviour that no component test predicts. The unit of assurance is the composed behaviour someone must own — not the parts.

  3. 10

    Make failure survivable before making it unlikely. Reversibility, containment, rollback and human override are assurance properties of the first rank. In agentic systems, the blast radius you designed is the strongest claim you have.

Accountability

  1. 11

    Retire evidence you no longer trust as deliberately as you retire code. Evidence ages: benchmarks saturate, evals drift from reality, yesterday's passing run describes yesterday's model. An assurance case built on stale evidence is manufactured certainty with a timestamp.

  2. 12

    Assurance is complete when a named person can honestly answer: do we know enough to ship — and how would we know if we were wrong? Not a committee, not a dashboard, not an agent. A person, with the evidence chain in front of them and the authority to say no.

The operating rule

Every assurance artefact carries a confidence question.

Every confidence claim carries an accountable decision.

Every accountable decision leaves an evidence answer.

A coverage area is run — is this justified confidence or performative coverage, and what evidence would change our mind? The 0–2 rubric produces scores — report the distribution with its interval, not a pass rate. A guardrail “passes” — the provider's claim is a hypothesis until independently probed. An agent acts — at which autonomy tier, under which decision, answered by what evidence?

The rule is a loop, not a checklist. A question interrogates the artefact; a claim survives only by forcing a decision — ship, hold, retest, retire; and a decision is discharged by an evidence answer: evidence designed, built and maintained to answer exactly that decision. Because every evidence answer is itself a new assurance artefact, it carries the next confidence question.

The through-line

Working software, adaptive outcomes, and calibrated evidence are the same commitment seen from three angles. Assurance is the thread that ties them to a name.

AI-Assurance (Governance) v1.0 · effective 2026-08-09 onwards