Building 11 min read

From evidence base to delivery: a production AI methodology

How we delivered 34 evidence-anchored AI briefings to a WA peer-advisory chapter: fact-checked literature review, multi-agent verification, one method.

In June 2026 we delivered 34 individual AI opportunity briefings to a 34-member peer-advisory chapter in Western Australia. Each member runs a different business, in a different industry, with a different operational shape, from solo principals to thirty-plus-staff firms. Each briefing had to be specific enough to be useful to that one operator and defensible enough to survive scrutiny from a professional reading about their own field.

This is a walk through how that work was produced. It is not a story about a clever prompt. It is a story about a pipeline: an evidence base built first and fact-checked openly, per-industry capabilities distilled from it, 34 briefings drafted against that evidence, a multi-agent review pass that caught five errors a human review had missed, and a chapter-level synthesis layered on top. The whole thing ran on Claude end to end, and the discipline at each stage is the point.

The challenge

Delivering substantive AI advice to 34 small businesses at once is harder than it sounds, and the difficulty is not volume. The difficulty is that a briefing for a lawyer and a briefing for a solar installer share almost nothing except the standard they have to meet. The lawyer needs to know what the courts now require of practitioners using AI, and where the published evidence on legal AI actually sits. The installer needs to know whether roof-modelling from an address is real yet, and how accurate. Get either one wrong and the reader, who is an expert in their own field, stops trusting the document on the first page.

So the requirement was uncomfortable: 34 documents, each anchored to evidence the reader could check, each tuned to a regulated or trades industry with its own frontier, and each written so the reader’s reaction was “I want this person on my side as my industry shifts,” not “I am being sold something.” Holding that posture across 34 briefings, consistently, is a methodology problem before it is a writing problem.

There was a second constraint that raised the bar further. More than a third of the room earns its living inside a regulated, document-bound profession, where a wrong claim is not just embarrassing but potentially a compliance problem. A briefing that overstated what AI can safely do in legal drafting, or understated a practitioner’s obligation to verify what a model produces, would be worse than useless to the reader. The advice had to be correct about the technology and correct about the rules governing its use, at the same time, for a dozen different regulatory regimes.

The six-stage pipeline

Stage 1: The frontier evidence base

Nothing in the briefings was allowed to rest on opinion. Underneath all of it sits a 13,000-word literature review on the state of applied AI in mid-2026, organised around ten capability categories and drawing on more than a hundred cited sources.

The review was drafted using Claude Fable 5, then subjected to three independent fact-check passes using Claude Opus 4.8. Across those passes, 135 individual fact-check findings were logged and resolved. The full corrections log is published as an appendix in the final document, and the pre-fact-check draft is preserved in the archive, so a reader can see exactly what changed and why. The published review is available in full at perthaiconsulting.com.au/resources/state-of-ai-mid-2026-literature-review-v1.0.pdf.

This is the load-bearing stage, not the preamble. The evidence layer of any responsible AI consulting work is the literature it rests on. Without it, the briefings would have been 34 sets of informed opinions. With it, every claim in every briefing traces back to a source the reader can inspect, and the correction trail behind that source is visible rather than hidden. That is the standard a regulated professional should expect from advice about their own practice, and it is the standard the rest of the pipeline inherits.

Stage 2: Industry-mapped capability surfacing

From the review, we distilled per-industry frontier capabilities, so different members received evidence anchored to their own context rather than a generic survey. Same source, different lens.

Legal members got the published evidence on legal AI: the Stanford RegLab measurements of hallucination in legal tools, the Federal Court’s GPN-AI verification obligations, the parallel state practice notes. Property-advisory members got the accuracy distribution of automated valuation models and the RICS responsible-AI standard that now governs regulated surveyors. Trades members got the real maturity of voice agents, around half of routine calls handled in production rather than the figure vendors quote. Signage and print members got the late-2025 shift in image models that finally render legible in-image text. Accounting-side members got Xero’s JAX rebuild and the honest state of bookkeeping AI. Energy-side members got roof-plane detection accuracy from address data. The clinical members got the AHPRA and TGA position on AI scribes as regulated devices. Each industry got its own slice of the same evidence base, which is why the briefings read as written for the reader rather than at them.

Stage 3: Personalised briefing drafting

The 34 briefings were drafted in Claude Cowork, each anchored to two things: the relevant evidence from the review, and the member’s own public footprint.

Every briefing followed the same six-section structure: an opening observation, where AI sits in that industry, what that means for the specific business, three numbered moves worth making, what to skip, and a closing line that holds a restrained posture rather than a pitch. The consistency mattered. Thirty-four documents that each found their own shape would have read as 34 different authors. One structure, held across all of them, reads as a single method applied 34 times.

The discipline at this stage was as much about what stayed off the page as what went on it. The briefings carried no internal product names, no catalogue phrasing, and no em dashes. The posture was intelligence-brief throughout: the work of someone who has read the field and is showing the reader what is now possible, not someone angling for a sale. That posture is easy to state and hard to hold across 34 drafts, which is exactly why the next stage exists.

Stage 4: Multi-agent quality assurance before delivery

Before anything was delivered, every briefing passed through a multi-agent Workflow on Claude. Seven review agents ran in parallel, each reading batches of briefings against a structured set of tone and accuracy criteria: intelligence-brief posture, evidence specificity, peer voice, the presence of any em dashes, catalogue language, internal product-name leakage, and hype vocabulary. A single synthesis agent then consolidated their findings into one report.

Eight agents in total, 240,000 tokens consumed, 51 tool uses, completed in four minutes. The verdict was that 30 of the 34 briefings read as pure intelligence-brief, four carried slight drift, and none crossed into a pitch.

The value showed up in five specific catches that the human review had missed:

  • Protected health information that had carried over from a clinical template into one of the legal-side briefings
  • An overclaim about the scope of AUSTRAC’s Tranche 2 obligations on an accounting-side briefing
  • A superlative (“single highest-leverage AI move”) on a property-advisory briefing, where the evidence did not support a ranking that strong
  • Catalogue vocabulary (“productise into a knowledge engine”) that had slipped into a financial-services briefing
  • Solicitation drift in the closing line of an advisory-services briefing

All five were fixed before delivery. This is the stage where Claude as multi-agent infrastructure earns its place in the pipeline. A single reviewer, however careful, reads linearly and tires. Seven agents reading against explicit criteria in parallel, with a synthesis pass on top, caught a privacy leak and a regulatory overclaim that a human had read past. Those are not stylistic notes. The PHI carryover alone would have been a serious problem had it reached a reader.

Stage 5: Chapter-level synthesis

The 34 briefings answered the question “what does AI mean for my business.” A second artifact answered a question no single member could: what does AI mean for this room as a whole.

We produced a chapter state report, roughly 2,800 words, generated via another multi-agent Workflow layered on top of all 34 briefings. It covered the composition of the chapter, the themes that recur across multiple members, the frontier capabilities that travel across industries, the regulatory dates landing on several members at once, and the referral loops that compound when members of a network adopt AI well together.

The synthesis surfaced patterns invisible from inside any one practice. That document throughput is the revenue ceiling for twelve of the members, all of whom sell a written document. That a single regulatory date, AUSTRAC’s Tranche 2 on 1 July 2026, lands on at least five members on the same morning, and that the most regulated members face two or three overlapping obligations at once. It also mapped the referral loops that form when one member’s AI adoption naturally pulls another member of the same network into the room: an accounting practice systemising its client knowledge becomes an infrastructure conversation; a valuation practice rebuilding its reporting needs a web presence layer. Each loop compounds, because each member’s work makes the next member’s referral warmer.

None of that is visible from one chair in the room. It is only visible once the 34 individual records are read as a single body of material, which is precisely what a synthesis layer is for, and it is the same reason a business with a working memory layer can see things across its own history that no single employee can.

Stage 6: Delivery and outcomes

The 34 briefings were delivered at a chapter meeting on 18 June 2026. The introduction took 45 seconds. The gift drew a round of applause in the meeting that day, before a single member had read their own brief.

Within a few days, the first commercial conversion conversation was underway.

The Claude toolchain

The engagement ran on Claude end to end, with the right tier matched to each job:

  • Claude Fable 5 for long-form drafting, including the literature review.
  • Claude Opus 4.8 for fact-checking the review and for the validation workflows.
  • Claude Sonnet 4.6 for production inference across the briefing work.
  • Claude Cowork for document generation and design rendering.
  • Claude Code for the engineering around the pipeline.
  • Claude Workflows for the multi-agent orchestration at the verification and synthesis stages.

The point of naming these is not the tooling itself. It is that a single vendor’s stack covered drafting, fact-checking, production inference, document rendering, engineering, and multi-agent orchestration, without stitching together a chain of disconnected services, and that different jobs went to different tiers deliberately rather than defaulting everything to one model.

What the methodology proves about the architecture

Underneath this engagement is the same six-function pattern we build into every client system: a knowledge layer (the literature review), an entity model (the 34 members and their industries), interaction capture (each member’s public footprint and briefing), memory (the consolidated record the synthesis drew on), a safety and egress layer (the multi-agent verification pass before delivery), and compounding outputs (the briefings, and the chapter report built on top of them).

The briefing pipeline is that architecture in miniature, run once over a chapter rather than continuously over a business. The literature review is a knowledge layer with explicit source rules. The verification Workflow is a safety and egress layer that checks output against evidence before it reaches a reader. The chapter report is a compounding output that feeds on everything captured below it. The same discipline that makes a clinical or legal AI system hold together in production is what made 34 briefings defensible. The shape of the work changes by domain. The architecture does not.

Built once, drawn upon every time

It is worth being plain about how this engagement relates to the next one. The literature review at the foundation of this work was not built for this work. It was built once, drawn upon here, and is drawn upon every time we do client work. The same is true of the six-function architecture, the multi-agent verification pattern, and the cowork prompts that hold the intelligence-brief posture across draft after draft. None of these were produced fresh for this chapter. Each one is part of the knowledge base we operate from, periodically updatable, and each engagement deepens it rather than using it up.

The practical consequence is that the marginal cost of the next engagement is lower than the last, the evidence base grows on every job, and nothing produced inside a client engagement is discarded once that engagement closes. The architecture is the asset, and the work compounds against it.

Outcomes and what comes next

The first commercial conversion is underway, a pipeline of follow-on conversations is forming, and the chapter report waits as the next artifact for the same audience. More durably, this pipeline is becoming the standard pattern across our delivery work: evidence base first, fact-checked in the open; per-context capability mapping; drafting against that evidence; a multi-agent verification pass before anything ships; and a synthesis layer that reads the whole body of work for patterns no single piece can show.

What makes the method repeatable is that none of it depends on the specific industries in that room. The same five stages would produce defensible briefings for a different chapter, a different profession, or a different evidence base, because the discipline lives in the architecture rather than the subject matter.

For the same shape at the scale of a single business, one practice, one finding, see What a good AI audit actually delivers.

A closing note

The interesting part of this engagement was never the writing. It was the chain from claim to source running through every stage, and the verification pass that made the difference between a document that reads well and a document that is safe to put in a regulated professional’s hands. That chain is the conversation we have with clients: not which model is cleverest this quarter, but what evidence the work rests on, what gets checked before anything leaves the building, and what compounds once it has.

Published 23 June 2026

Perth AI Consulting delivers AI opportunity analysis for small and medium businesses. Start with a conversation.

Prepared by Claude, directed and approved by PAC.

More from Thinking

Building 7 min read

Why we let AI run the interviews (and why we never let it pretend to be human)

AI-conducted interviews compress weeks of stakeholder discovery into days, standardise what gets asked, and lower the guard that distorts honest answers.

Adoption 14 min read

How AI capability actually moves through a business

The decisive variable in SME AI adoption is the human absorption sequence, not the tooling. A working framework from observation across WA businesses.

Evaluation 7 min read

AHPRA advertising rules for psychologist websites

Recovery stories, 'specialist', 'clinical psychologist', and endorsement titles are where psychology sites breach the National Law. A practical read-through.

Adoption 6 min read

Customer service AI has finally grown up

Chatbots and AI receptionists earned their bad reputation. What changed, why the trick is in the data, and how the mature version answers every call without replacing anyone.

Evaluation 6 min read

Who can use the titles 'Dr', 'Specialist', and 'Surgeon'?

AHPRA restricts 'specialist' and 'surgeon' to specific registrations, and 'Dr' has its own rule. What health practice websites can and cannot claim.

Adoption 5 min read

Your best people hate writing reports

The operators you promote are brilliant at the work and allergic to reporting. A scheduled AI call interviews them, drafts the briefing, and they approve it. No ego, no politics, no blank page.

Building 6 min read

Your website isn't just for humans anymore

How to build a chatbot that keeps itself up to date, can't leak client information, and won't answer beyond what you've published. The answer was sitting in plain sight.

Evaluation 7 min read

Can you show Google reviews on your health practice website?

AHPRA bans clinical testimonials, even true ones, but service reviews are fine. What that means for the Google reviews widget on your practice site.

Evaluation 7 min read

What AHPRA's advertising rules mean for your website

Your practice website is advertising under the National Law. What AHPRA's rules prohibit, who is responsible, and how to check your own site.

Evaluation 8 min read

Is it safe to paste client data into ChatGPT?

Short answer: it depends on one setting, and most people have it wrong. What ChatGPT, Claude and Copilot do with your data, and what the Privacy Act expects.

Evaluation 6 min read

What a good AI audit actually delivers

The audit report named one recommendation specific enough to check. What the Build engagement that followed looked like, shown through one real engagement, generalised.

Evaluation 7 min read

AI and video, Mid-2026: the models can watch now, not just listen

AI could always transcribe video. It can now read the frames as well, and every hour of footage a business owns becomes something it can question.

Technical 5 min read

Why the privacy case against cloud AI memory isn't paranoia

An AI knowledge base concentrates everything sensitive a business holds. Dated 2026 incidents show what cloud custody means once legal process gets involved.

Technical 5 min read

Your AI knowledge base is an attack surface

A knowledge base an AI agent can read and write is a productivity tool, and dated 2026 incidents show it is also somewhere an attacker can plant instructions.

Adoption 5 min read

The real asset in an AI knowledge base isn't the notes

In every AI-maintained knowledge base, one file carries the owner's judgement and compounds. The wiki pages are the least valuable part.

Evaluation 5 min read

The missing measurement in the AI second-brain boom

Every claim about AI knowledge bases saving time is self-reported. The one controlled experiment measured token economics, not benefit.

Technical 9 min read

The six functions of a working AI system

A working AI system is six functions doing six jobs. When all six connect, hallucinations get caught, outputs hold steady, and models become swappable.

Technical 7 min read

Supervised autonomy: the middle path for AI architecture

Between drafts you approve and agents you hope about sits the middle path: an envelope of authorised routine work, supervised, audited, and yours to widen.

Evaluation 5 min read

The state of applied AI in Mid-2026

Our literature review of applied AI in mid-2026: ten capability categories, three fact-check passes, written for operational leaders.

Evaluation 8 min read

AI in building inspections, Mid-2026

AI defect detection is strong on obvious defects and weak on the subtle ones where liability lives. Which capabilities fit inspection work in 2026.

Evaluation 8 min read

AI in property valuation, Mid-2026

AVMs are reliable enough for triage, not for the final word on contested property. What has shifted in valuation work by mid-2026, and what has not.

Evaluation 8 min read

AI in family law, Mid-2026

Federal Court practice note GPN-AI makes AI verification a professional obligation. What the courts now require, and what the evidence says about legal AI.

Technical 9 min read

How to design a PHI redaction system for clinical AI

PHI redaction is part of a clinical AI tool's architecture, not a feature you add. What the literature says it should look like, and how we built it.

Building 9 min read

How we built on-device de-identification so AI never sees real names

Most AI privacy is a policy. Ours is architecture: an NER model runs in the browser and strips names before anything leaves the device.

Technical 7 min read

Your agency's clients are about to ask why this costs so much

A solo consultant built in three weeks what your agency quoted twelve for. The client doesn't know why yet. The agencies that survive change what they sell.

Adoption 6 min read

What do you love doing? What do you hate doing?

Ask people what they love doing and what they hate doing, then show them AI is coming for the second list. Why the reframe works, and how it fails.

Technical 7 min read

Why I don't use n8n (and what I do instead)

n8n demos well. But a compelling demo and a reliable production system are different things, and the distance between them is where businesses get hurt.

Technical 10 min read

Your codebase was not built for AI. That's the actual problem.

Amazon's mandatory meeting about AI breaking production is an architecture story: codebases built for human maintainers only, now maintained by AI.

Adoption 4 min read

Your team has AI licences. You don't have an AI system.

Fifteen people, fifteen separate AI accounts, no shared context. The problem isn't the tool; it's the architecture around it. Here's the fix.

Building 7 min read

Your $2,000 day starts the night before: our system keeps you on the tools, not on the phone

Optimised routes overnight, automatic customer notifications, and promises the system keeps or corrects. A scheduling system that protects your daily rate.

Evaluation 4 min read

The fastest way for an executive to get across AI

AI moves faster than any executive can track. One focused conversation, one written report, and a decision you can act on: your time stays on the business.

Building 6 min read

Your IT department will take 18 months. You need this working by next quarter.

Senior leaders know what they need built; the gap is time. A prototype gets the tool working now and hands IT a validated blueprint for later.

Adoption 4 min read

What if you had perfect memory across every client?

Every practice captures more than it can recall. AI gives practitioners perfect memory across every client, so preparation becomes thinking time.

Building 8 min read

We built an AI invoice verifier. Here's where it hits a wall.

We built an AI invoice verifier and watched a fake beat a real invoice. Why document analysis alone cannot stop fraud, and the five layers that can.

Building 5 min read

How to build an AI chatbot that doesn't lie to your customers

Woolworths scripted its AI to talk about its mother. The business fix is honesty; the technical fix is architecture that prevents fabrication by design.

Technical 9 min read

Why AI safety features are load-bearing architecture, not political decoration

The 'woke AI' label came from real failures, but they were engineering failures, not safety failures. The difference matters wherever errors have consequences.

Adoption 3 min read

Woolworths' AI told a customer it had a mother. That's a problem.

Woolworths' AI assistant Olive was scripted to talk about its mother and uncle. When callers realised, trust broke instantly. The fix is honesty.

Evaluation 5 min read

Google is no longer the only way your customers find you

Customers now find businesses through ChatGPT, Perplexity, and Gemini. The sites AI cites are structured differently to the sites Google ranks.

Evaluation 4 min read

Two types of AI audit: and how to know which one you need

Where do we start with AI? It depends on whether you need to find the opportunities or reclaim the time. Two audits, two perspectives, one goal.

Evaluation 4 min read

The personal workflow analysis: what watching a real workday reveals about automation

People describe the work they value, not the work that eats their time. Recording a real workday reveals the automation opportunities interviews miss.

Evaluation 4 min read

AI audit that starts with your business

An operations-first AI audit starts with how your business actually runs, and only recommends AI where the evidence says it will work.

Building 6 min read

What production AI teaches you that demos never will

The gap between a demo and a working system is where the useful lessons live. Architecture, framing, privacy, adoption: the patterns repeat every time.

Adoption 6 min read

The psychology of why your team won't use AI

You buy the tool, run the demo, and three months later nobody is using it. Five predictable psychological barriers, each with a strategy that works.

Technical 4 min read

Stop telling AI what NOT to do (and what to say instead)

Instructions built on prohibitions make AI cautious and generic. Describing what you want instead transforms the output, and the reason comes from psychology.

Building 5 min read

How we turned generic AI into a specialist: and what that means for your business

Mediocre AI output is rarely the model's fault. Three structural changes that turn the same model from generic to specialist-grade.

Evaluation 5 min read

Your business has 9 customer touchpoints. AI can fix the 6 you're dropping.

You pay to get customers to your door, then lose them to missed follow-up. AI can handle the six touchpoints most businesses drop.

Technical 5 min read

What happens to your data when you press 'Send' on an AI tool

Businesses send customer data to AI tools without knowing what happens during processing. The spectrum of AI privacy is wider than you think.