From evidence base to delivery: a production AI methodology
How we delivered 34 evidence-anchored AI briefings to a WA peer-advisory chapter: fact-checked literature review, multi-agent verification, one method.
In June 2026 we delivered 34 individual AI opportunity briefings to a 34-member peer-advisory chapter in Western Australia. Each member runs a different business, in a different industry, with a different operational shape, from solo principals to thirty-plus-staff firms. Each briefing had to be specific enough to be useful to that one operator and defensible enough to survive scrutiny from a professional reading about their own field.
This is a walk through how that work was produced. It is not a story about a clever prompt. It is a story about a pipeline: an evidence base built first and fact-checked openly, per-industry capabilities distilled from it, 34 briefings drafted against that evidence, a multi-agent review pass that caught five errors a human review had missed, and a chapter-level synthesis layered on top. The whole thing ran on Claude end to end, and the discipline at each stage is the point.
The challenge
Delivering substantive AI advice to 34 small businesses at once is harder than it sounds, and the difficulty is not volume. The difficulty is that a briefing for a lawyer and a briefing for a solar installer share almost nothing except the standard they have to meet. The lawyer needs to know what the courts now require of practitioners using AI, and where the published evidence on legal AI actually sits. The installer needs to know whether roof-modelling from an address is real yet, and how accurate. Get either one wrong and the reader, who is an expert in their own field, stops trusting the document on the first page.
So the requirement was uncomfortable: 34 documents, each anchored to evidence the reader could check, each tuned to a regulated or trades industry with its own frontier, and each written so the reader’s reaction was “I want this person on my side as my industry shifts,” not “I am being sold something.” Holding that posture across 34 briefings, consistently, is a methodology problem before it is a writing problem.
There was a second constraint that raised the bar further. More than a third of the room earns its living inside a regulated, document-bound profession, where a wrong claim is not just embarrassing but potentially a compliance problem. A briefing that overstated what AI can safely do in legal drafting, or understated a practitioner’s obligation to verify what a model produces, would be worse than useless to the reader. The advice had to be correct about the technology and correct about the rules governing its use, at the same time, for a dozen different regulatory regimes.
The six-stage pipeline
Stage 1: The frontier evidence base
Nothing in the briefings was allowed to rest on opinion. Underneath all of it sits a 13,000-word literature review on the state of applied AI in mid-2026, organised around ten capability categories and drawing on more than a hundred cited sources.
The review was drafted using Claude Fable 5, then subjected to three independent fact-check passes using Claude Opus 4.8. Across those passes, 135 individual fact-check findings were logged and resolved. The full corrections log is published as an appendix in the final document, and the pre-fact-check draft is preserved in the archive, so a reader can see exactly what changed and why. The published review is available in full at perthaiconsulting.com.au/resources/state-of-ai-mid-2026-literature-review-v1.0.pdf.
This is the load-bearing stage, not the preamble. The evidence layer of any responsible AI consulting work is the literature it rests on. Without it, the briefings would have been 34 sets of informed opinions. With it, every claim in every briefing traces back to a source the reader can inspect, and the correction trail behind that source is visible rather than hidden. That is the standard a regulated professional should expect from advice about their own practice, and it is the standard the rest of the pipeline inherits.
Stage 2: Industry-mapped capability surfacing
From the review, we distilled per-industry frontier capabilities, so different members received evidence anchored to their own context rather than a generic survey. Same source, different lens.
Legal members got the published evidence on legal AI: the Stanford RegLab measurements of hallucination in legal tools, the Federal Court’s GPN-AI verification obligations, the parallel state practice notes. Property-advisory members got the accuracy distribution of automated valuation models and the RICS responsible-AI standard that now governs regulated surveyors. Trades members got the real maturity of voice agents, around half of routine calls handled in production rather than the figure vendors quote. Signage and print members got the late-2025 shift in image models that finally render legible in-image text. Accounting-side members got Xero’s JAX rebuild and the honest state of bookkeeping AI. Energy-side members got roof-plane detection accuracy from address data. The clinical members got the AHPRA and TGA position on AI scribes as regulated devices. Each industry got its own slice of the same evidence base, which is why the briefings read as written for the reader rather than at them.
Stage 3: Personalised briefing drafting
The 34 briefings were drafted in Claude Cowork, each anchored to two things: the relevant evidence from the review, and the member’s own public footprint.
Every briefing followed the same six-section structure: an opening observation, where AI sits in that industry, what that means for the specific business, three numbered moves worth making, what to skip, and a closing line that holds a restrained posture rather than a pitch. The consistency mattered. Thirty-four documents that each found their own shape would have read as 34 different authors. One structure, held across all of them, reads as a single method applied 34 times.
The discipline at this stage was as much about what stayed off the page as what went on it. The briefings carried no internal product names, no catalogue phrasing, and no em dashes. The posture was intelligence-brief throughout: the work of someone who has read the field and is showing the reader what is now possible, not someone angling for a sale. That posture is easy to state and hard to hold across 34 drafts, which is exactly why the next stage exists.
Stage 4: Multi-agent quality assurance before delivery
Before anything was delivered, every briefing passed through a multi-agent Workflow on Claude. Seven review agents ran in parallel, each reading batches of briefings against a structured set of tone and accuracy criteria: intelligence-brief posture, evidence specificity, peer voice, the presence of any em dashes, catalogue language, internal product-name leakage, and hype vocabulary. A single synthesis agent then consolidated their findings into one report.
Eight agents in total, 240,000 tokens consumed, 51 tool uses, completed in four minutes. The verdict was that 30 of the 34 briefings read as pure intelligence-brief, four carried slight drift, and none crossed into a pitch.
The value showed up in five specific catches that the human review had missed:
- Protected health information that had carried over from a clinical template into one of the legal-side briefings
- An overclaim about the scope of AUSTRAC’s Tranche 2 obligations on an accounting-side briefing
- A superlative (“single highest-leverage AI move”) on a property-advisory briefing, where the evidence did not support a ranking that strong
- Catalogue vocabulary (“productise into a knowledge engine”) that had slipped into a financial-services briefing
- Solicitation drift in the closing line of an advisory-services briefing
All five were fixed before delivery. This is the stage where Claude as multi-agent infrastructure earns its place in the pipeline. A single reviewer, however careful, reads linearly and tires. Seven agents reading against explicit criteria in parallel, with a synthesis pass on top, caught a privacy leak and a regulatory overclaim that a human had read past. Those are not stylistic notes. The PHI carryover alone would have been a serious problem had it reached a reader.
Stage 5: Chapter-level synthesis
The 34 briefings answered the question “what does AI mean for my business.” A second artifact answered a question no single member could: what does AI mean for this room as a whole.
We produced a chapter state report, roughly 2,800 words, generated via another multi-agent Workflow layered on top of all 34 briefings. It covered the composition of the chapter, the themes that recur across multiple members, the frontier capabilities that travel across industries, the regulatory dates landing on several members at once, and the referral loops that compound when members of a network adopt AI well together.
The synthesis surfaced patterns invisible from inside any one practice. That document throughput is the revenue ceiling for twelve of the members, all of whom sell a written document. That a single regulatory date, AUSTRAC’s Tranche 2 on 1 July 2026, lands on at least five members on the same morning, and that the most regulated members face two or three overlapping obligations at once. It also mapped the referral loops that form when one member’s AI adoption naturally pulls another member of the same network into the room: an accounting practice systemising its client knowledge becomes an infrastructure conversation; a valuation practice rebuilding its reporting needs a web presence layer. Each loop compounds, because each member’s work makes the next member’s referral warmer.
None of that is visible from one chair in the room. It is only visible once the 34 individual records are read as a single body of material, which is precisely what a synthesis layer is for, and it is the same reason a business with a working memory layer can see things across its own history that no single employee can.
Stage 6: Delivery and outcomes
The 34 briefings were delivered at a chapter meeting on 18 June 2026. The introduction took 45 seconds. The gift drew a round of applause in the meeting that day, before a single member had read their own brief.
Within a few days, the first commercial conversion conversation was underway.
The Claude toolchain
The engagement ran on Claude end to end, with the right tier matched to each job:
- Claude Fable 5 for long-form drafting, including the literature review.
- Claude Opus 4.8 for fact-checking the review and for the validation workflows.
- Claude Sonnet 4.6 for production inference across the briefing work.
- Claude Cowork for document generation and design rendering.
- Claude Code for the engineering around the pipeline.
- Claude Workflows for the multi-agent orchestration at the verification and synthesis stages.
The point of naming these is not the tooling itself. It is that a single vendor’s stack covered drafting, fact-checking, production inference, document rendering, engineering, and multi-agent orchestration, without stitching together a chain of disconnected services, and that different jobs went to different tiers deliberately rather than defaulting everything to one model.
What the methodology proves about the architecture
Underneath this engagement is the same six-function pattern we build into every client system: a knowledge layer (the literature review), an entity model (the 34 members and their industries), interaction capture (each member’s public footprint and briefing), memory (the consolidated record the synthesis drew on), a safety and egress layer (the multi-agent verification pass before delivery), and compounding outputs (the briefings, and the chapter report built on top of them).
The briefing pipeline is that architecture in miniature, run once over a chapter rather than continuously over a business. The literature review is a knowledge layer with explicit source rules. The verification Workflow is a safety and egress layer that checks output against evidence before it reaches a reader. The chapter report is a compounding output that feeds on everything captured below it. The same discipline that makes a clinical or legal AI system hold together in production is what made 34 briefings defensible. The shape of the work changes by domain. The architecture does not.
Built once, drawn upon every time
It is worth being plain about how this engagement relates to the next one. The literature review at the foundation of this work was not built for this work. It was built once, drawn upon here, and is drawn upon every time we do client work. The same is true of the six-function architecture, the multi-agent verification pattern, and the cowork prompts that hold the intelligence-brief posture across draft after draft. None of these were produced fresh for this chapter. Each one is part of the knowledge base we operate from, periodically updatable, and each engagement deepens it rather than using it up.
The practical consequence is that the marginal cost of the next engagement is lower than the last, the evidence base grows on every job, and nothing produced inside a client engagement is discarded once that engagement closes. The architecture is the asset, and the work compounds against it.
Outcomes and what comes next
The first commercial conversion is underway, a pipeline of follow-on conversations is forming, and the chapter report waits as the next artifact for the same audience. More durably, this pipeline is becoming the standard pattern across our delivery work: evidence base first, fact-checked in the open; per-context capability mapping; drafting against that evidence; a multi-agent verification pass before anything ships; and a synthesis layer that reads the whole body of work for patterns no single piece can show.
What makes the method repeatable is that none of it depends on the specific industries in that room. The same five stages would produce defensible briefings for a different chapter, a different profession, or a different evidence base, because the discipline lives in the architecture rather than the subject matter.
For the same shape at the scale of a single business, one practice, one finding, see What a good AI audit actually delivers.
A closing note
The interesting part of this engagement was never the writing. It was the chain from claim to source running through every stage, and the verification pass that made the difference between a document that reads well and a document that is safe to put in a regulated professional’s hands. That chain is the conversation we have with clients: not which model is cleverest this quarter, but what evidence the work rests on, what gets checked before anything leaves the building, and what compounds once it has.