The Loom · Toolkit 2.0.0 released
Find the right problem.
Ship it under control.
Two harnesses turn an ambiguous mandate into audit-ready software. Run and Operations then return evidence to Discovery, so the institution learns from what the software actually does.
Executive view
A governed route from ambiguous mandate to accountable software.
The Loom is for institutions that want the speed of AI-assisted delivery without giving an agent authority over scope, controls or release.
Start before code
Frame the outcome, evidence, boundaries and decision rights before implementation begins.
Regulated leaders
Built for sponsors, product owners, risk leaders and delivery teams working under real institutional constraints.
One closed loop
Discovery, delivery and operations share evidence instead of handing work across disconnected phases.
Toolkit 2.0.0
The current public release packages the method for repository adoption.
Reference-build validated
Exercised end to end while building the synthetic Open Finance Backoffice reference portal.
No customer production use
It has not been used to deliver or operate a customer production system.
The idea
The metaphor is exact.
Every part of a physical loom maps to a durable part of the operating method.
Always-on principles
Evidence, bounded work, human authority, quality and traceability remain under tension throughout.
Discovery + delivery
One harness finds the right problem. The other ships the chosen solution under control.
AI agents
Specialised agents move through the harness to research, build, review, test and evidence.
The context brain
Mandate constraints, solution-domain knowledge and institutional context determine what gets woven.
Audit-ready software
The output is a working solution with its controls, lineage and release evidence constructed into the line.
Mandate to outcome
One traceable control chain—not a collection of AI tools.
The same success measures that justify the problem become the evaluation targets for the released product. Every step leaves an owned artifact.
Mandate
Leaders name the outcome, risk posture, constraints and accountable owners.
Problem
Discovery tests the framing and produces one gate-green, evidenced hand-off.
Solution
Product and architecture assurance shape the direction before autonomous delivery.
Release
Quality gates, protected controls and accountable approval govern promotion.
Outcome
A fresh product evaluation scores the release against Discovery’s success measures.
Signal
Incidents, drift, regulation and customer evidence are triaged back into the loop.
The two harnesses
A double diamond: find the right problem, then deliver it.
The diamonds meet at one enforced waist: a gate-green hand-off. Discovery may stop a weak problem early; delivery evidence may legitimately send the work back.
Deploy → observe → triage. Most signals become a delivery fix or risk-register update. Only evidence that challenges the original framing reopens Discovery.
↶ Feedback edge to D2 evidenceDiamond 01
Discovery: evidence in, problem out.
Before code, every claim, boundary, data-risk decision, prototype and stakeholder reaction becomes a traceable artifact.
A disposable wireframe tests the framing. It never binds the production solution.
Framing
A falsifiable problem, target user and success measure.
Evidence
Claims trace to logged signals rather than opinion.
Scope
Stakeholders and in/out boundaries are named.
No solutioning
Discovery does not leak into production design.
Synthesis
Themes trace to signals and prioritisation is explicit.
Data governance
Risks and mandate-specific drivers resolve against the register.
Brand
Stakeholder artifacts render through the mounted brand profile.
Tangibility
A disposable, low-fidelity prototype makes the direction visible.
Validation
Named stakeholders react and the evidence loop closes.
Diamond 02
Delivery: an autonomous loop that proposes, never disposes.
One loop takes one eligible item end to end. The agent authors and verifies; a protected control plane keeps approval accountable—per change by default, or through a narrowly governed routine envelope.
- 01
Pick
Take the first eligible item. The waist gate rejects untraced features.
- 02
Isolate
One story, one branch, one session; concurrent work uses isolated worktrees.
- 03
Specify
Contract, generated types and acceptance tests precede implementation.
- 04
Implement
Build to green with synthetic data, audit and lineage in the same change.
- 05
Review
Bounded reviewers judge hard stops and contract conformance.
- 06
Human approval
A protected control plane requires accountable approval; regulated work can require formal four-eyes.
- 07
Deploy
Promotion, live smoke checks and rollback remain governed.
- 08
Evidence
Release gates rerun at the shipped commit and seal the evidence bundle.
The pattern
The context brain is why the method compounds.
A generic harness becomes institutional capability when context is governed as an owned, versioned and auditable asset.
Mandate context
Commercial goals, regulation where applicable, deadlines, risk posture and controls.
Solution domain
Contracts, data models, personas, scopes and synthetic datasets.
Institutional DNA
Approval routes, architecture, approved technology, terminology, brand and voice.
If a competitor copied the codebase tomorrow, it would still lack the accumulated, governed context that makes the software belong to the institution.
Continuous assurance · always-on
Agents do the recurring work. People retain judgement.
The assurance lifecycle re-runs on changes, schedules and regulatory events, keeping evidence current to the last trigger instead of the last meeting.
- 01
Watch
Regulatory and risk horizon scan
- 02
Assess
Risk and impact review
- 03
Check
Contract and hard-stop conformance
- 04
Test
Controls and quality gates
- 05
Evidence
Audit trail and lineage capture
- 06
Confirm
Agent-prepared reporting plus human four-eyes
Harness governance · who assures the AI?
AI proposes. Humans and a protected control plane dispose.
The regulated reference catalogue closes the gaps that let an ungoverned agent self-review, self-merge, deploy and edit its own guardrails.
No self-merge
Human-approved merges through enforced branch protection.
Immutable controls
Guardrails, gates and CI sit outside the agent's write scope.
Sealed evidence
Traceability is externally anchored and tamper-evident.
Least privilege
The agent uses a bounded identity and vaulted secrets.
Governed promotion
Staged release, a human production gate and rehearsed rollback.
Model risk
The harness itself is governed as an AI system.
Discovery first
A gate-green hand-off is the entry condition for delivery.
Mounted seams
Risk registers and brand profiles are mounted, not hard-coded.
Diverge, then decide
Delivery explores several directions before converging.
Cease use
A kill switch and named accountable officer are mandatory.
Residency control
Model traffic passes through governed gateways and DLP.
Derive, do not retrieve
A sealed runtime distinguishes reasoning from answer mining.
Graduated autonomy
A narrow routine-change lane can move approval from each change to a second-line-owned, expiring envelope. Approval is relocated, never removed.
Human determinations
Religious and ethical determinations are human-issued context, never agent work-product.
A control is not “bank-grade” merely because it is documented.
The Loom separates repository mechanics from platform enforcement and real operating evidence.
- 01Absent
No credible control exists.
- 02Defined
The decision and owner are documented.
- 03Mechanically validated
Repository machinery tests the claim.
- 04Platform enforced
The agent cannot bypass the control.
- 05Organisationally enforced
Operating evidence proves it works in practice.
Proof and limits
Validated on a reference build. Stated plainly.
The synthetic Open Finance Backoffice demonstrates a real gated system and a reusable method. It is not evidence of customer production use.
Reference-build validated, not production-proven
It has not been used to deliver or operate a customer production system. It has not cleared live production scale or a regulator examination.
One domain is early evidence
Legacy integration, real data and organisational change remain the true cost curve for other institutions.
Spend measured; value unproven
Token telemetry now measures delivery spend by iteration and milestone. The value half—and any implied ROI—remains unbuilt.
Comprehension debt remains
Decision logs make agent reasoning replayable, but they do not prove that human reviewers still understand a growing codebase.
The brain must be curated
Bad context compounds as quickly as good context, so ownership and quality control are part of the method.
AI model risk is real
Agents require validation, monitoring and independent challenge; automation does not remove accountability.
A bank-neutral, synthetic-only portal used to exercise the harness end to end—not a customer production deployment.
Read the build record →Applications of its evidence, specification and human-authority principles—not claims of full regulated-harness adoption.
Explore the portfolio →Adopt the method
Start with one real mandate and leave a reusable capability behind.
Mount the institution's controls and context, run one gated discovery, and deliver one bounded outcome with human accountability intact.