MiddleLeap
Menu

The Loom · Toolkit 2.0.0 released

Find the right problem.
Ship it under control.

Two harnesses turn an ambiguous mandate into audit-ready software. Run and Operations then return evidence to Discovery, so the institution learns from what the software actually does.

Mandate → outcome / closed loopTwo harnesses · one feedback arc
EvidenceBoundariesAuthorityQualityTraceability
Harness 01DiscoveryDiscover → Define · D1—D9
Harness 02DeliveryDevelop → Deliver · Q1—Q5
The patternMandate context brainConstraints · Domain · Institutional context
AI agents weave continuously
The clothAudit-ready software
The third arcRun / OperationsReality tests the framing

Executive view

A governed route from ambiguous mandate to accountable software.

The Loom is for institutions that want the speed of AI-assisted delivery without giving an agent authority over scope, controls or release.

01 / Mandate

Start before code

Frame the outcome, evidence, boundaries and decision rights before implementation begins.

02 / For whom

Regulated leaders

Built for sponsors, product owners, risk leaders and delivery teams working under real institutional constraints.

03 / Capability

One closed loop

Discovery, delivery and operations share evidence instead of handing work across disconnected phases.

04 / Release

Toolkit 2.0.0

The current public release packages the method for repository adoption.

05 / Evidence

Reference-build validated

Exercised end to end while building the synthetic Open Finance Backoffice reference portal.

06 / Boundary

No customer production use

It has not been used to deliver or operate a customer production system.

134 / ~139Stories to done in the synthetic reference build
2 + 1Harnesses and the Run feedback arc
100%Merges approved by people in the reference build
0Real customer records used

The idea

The metaphor is exact.

Every part of a physical loom maps to a durable part of the operating method.

Warp

Always-on principles

Evidence, bounded work, human authority, quality and traceability remain under tension throughout.

Harnesses

Discovery + delivery

One harness finds the right problem. The other ships the chosen solution under control.

Shuttle

AI agents

Specialised agents move through the harness to research, build, review, test and evidence.

Pattern

The context brain

Mandate constraints, solution-domain knowledge and institutional context determine what gets woven.

Cloth

Audit-ready software

The output is a working solution with its controls, lineage and release evidence constructed into the line.

Mandate to outcome

One traceable control chain—not a collection of AI tools.

The same success measures that justify the problem become the evaluation targets for the released product. Every step leaves an owned artifact.

01

Mandate

Leaders name the outcome, risk posture, constraints and accountable owners.

02

Problem

Discovery tests the framing and produces one gate-green, evidenced hand-off.

03

Solution

Product and architecture assurance shape the direction before autonomous delivery.

04

Release

Quality gates, protected controls and accountable approval govern promotion.

05

Outcome

A fresh product evaluation scores the release against Discovery’s success measures.

06

Signal

Incidents, drift, regulation and customer evidence are triaged back into the loop.

Signal routedThe next decision starts with evidence from the last one.

The two harnesses

A double diamond: find the right problem, then deliver it.

The diamonds meet at one enforced waist: a gate-green hand-off. Discovery may stop a weak problem early; delivery evidence may legitimately send the work back.

01DiscoverDiverge around evidence
02DefineConverge on one problem
Waist gateAgreed hand-off
03DevelopDiverge across solutions
04DeliverConverge under control
Run / Operations

Deploy → observe → triage. Most signals become a delivery fix or risk-register update. Only evidence that challenges the original framing reopens Discovery.

Diamond 01

Discovery: evidence in, problem out.

Before code, every claim, boundary, data-risk decision, prototype and stakeholder reaction becomes a traceable artifact.

Prototype boundaryBrand-real. Behaviour-hollow.

A disposable wireframe tests the framing. It never binds the production solution.

D1

Framing

A falsifiable problem, target user and success measure.

D2

Evidence

Claims trace to logged signals rather than opinion.

D3

Scope

Stakeholders and in/out boundaries are named.

D4

No solutioning

Discovery does not leak into production design.

D5

Synthesis

Themes trace to signals and prioritisation is explicit.

D6

Data governance

Risks and mandate-specific drivers resolve against the register.

D7

Brand

Stakeholder artifacts render through the mounted brand profile.

D8

Tangibility

A disposable, low-fidelity prototype makes the direction visible.

D9

Validation

Named stakeholders react and the evidence loop closes.

Diamond 02

Delivery: an autonomous loop that proposes, never disposes.

One loop takes one eligible item end to end. The agent authors and verifies; a protected control plane keeps approval accountable—per change by default, or through a narrowly governed routine envelope.

  1. 01

    Pick

    Take the first eligible item. The waist gate rejects untraced features.

  2. 02

    Isolate

    One story, one branch, one session; concurrent work uses isolated worktrees.

  3. 03

    Specify

    Contract, generated types and acceptance tests precede implementation.

  4. 04

    Implement

    Build to green with synthetic data, audit and lineage in the same change.

  5. 05

    Review

    Bounded reviewers judge hard stops and contract conformance.

  6. 06

    Human approval

    A protected control plane requires accountable approval; regulated work can require formal four-eyes.

  7. 07

    Deploy

    Promotion, live smoke checks and rollback remain governed.

  8. 08

    Evidence

    Release gates rerun at the shipped commit and seal the evidence bundle.

PII guard+Spec tripwire+Test-integrity tripwireControls apply at the moment of edit.

The pattern

The context brain is why the method compounds.

A generic harness becomes institutional capability when context is governed as an owned, versioned and auditable asset.

01What must be obeyed

Mandate context

Commercial goals, regulation where applicable, deadlines, risk posture and controls.

02What is being built

Solution domain

Contracts, data models, personas, scopes and synthetic datasets.

03How the entity works

Institutional DNA

Approval routes, architecture, approved technology, terminology, brand and voice.

The moat test

If a competitor copied the codebase tomorrow, it would still lack the accumulated, governed context that makes the software belong to the institution.

Continuous assurance · always-on

Agents do the recurring work. People retain judgement.

The assurance lifecycle re-runs on changes, schedules and regulatory events, keeping evidence current to the last trigger instead of the last meeting.

  1. 01

    Watch

    Regulatory and risk horizon scan

  2. 02

    Assess

    Risk and impact review

  3. 03

    Check

    Contract and hard-stop conformance

  4. 04

    Test

    Controls and quality gates

  5. 05

    Evidence

    Audit trail and lineage capture

  6. 06

    Confirm

    Agent-prepared reporting plus human four-eyes

Harness governance · who assures the AI?

AI proposes. Humans and a protected control plane dispose.

The regulated reference catalogue closes the gaps that let an ungoverned agent self-review, self-merge, deploy and edit its own guardrails.

HG-0001

No self-merge

Human-approved merges through enforced branch protection.

HG-0002

Immutable controls

Guardrails, gates and CI sit outside the agent's write scope.

HG-0003

Sealed evidence

Traceability is externally anchored and tamper-evident.

HG-0004

Least privilege

The agent uses a bounded identity and vaulted secrets.

HG-0005

Governed promotion

Staged release, a human production gate and rehearsed rollback.

HG-0006

Model risk

The harness itself is governed as an AI system.

HG-0007

Discovery first

A gate-green hand-off is the entry condition for delivery.

HG-0008

Mounted seams

Risk registers and brand profiles are mounted, not hard-coded.

HG-0009

Diverge, then decide

Delivery explores several directions before converging.

HG-0010

Cease use

A kill switch and named accountable officer are mandatory.

HG-0011

Residency control

Model traffic passes through governed gateways and DLP.

HG-0012

Derive, do not retrieve

A sealed runtime distinguishes reasoning from answer mining.

HG-0013

Graduated autonomy

A narrow routine-change lane can move approval from each change to a second-line-owned, expiring envelope. Approval is relocated, never removed.

HG-0014

Human determinations

Religious and ethical determinations are human-issued context, never agent work-product.

Honest self-grade

A control is not “bank-grade” merely because it is documented.

The Loom separates repository mechanics from platform enforcement and real operating evidence.

  1. 01Absent

    No credible control exists.

  2. 02Defined

    The decision and owner are documented.

  3. 03Mechanically validated

    Repository machinery tests the claim.

  4. 04Platform enforced

    The agent cannot bypass the control.

  5. 05Organisationally enforced

    Operating evidence proves it works in practice.

Proof and limits

Validated on a reference build. Stated plainly.

The synthetic Open Finance Backoffice demonstrates a real gated system and a reusable method. It is not evidence of customer production use.

Reference-build validated, not production-proven

It has not been used to deliver or operate a customer production system. It has not cleared live production scale or a regulator examination.

One domain is early evidence

Legacy integration, real data and organisational change remain the true cost curve for other institutions.

Spend measured; value unproven

Token telemetry now measures delivery spend by iteration and milestone. The value half—and any implied ROI—remains unbuilt.

Comprehension debt remains

Decision logs make agent reasoning replayable, but they do not prove that human reviewers still understand a growing codebase.

The brain must be curated

Bad context compounds as quickly as good context, so ownership and quality control are part of the method.

AI model risk is real

Agents require validation, monitoring and independent challenge; automation does not remove accountability.

Adopt the method

Start with one real mandate and leave a reusable capability behind.

Mount the institution's controls and context, run one gated discovery, and deliver one bounded outcome with human accountability intact.