Free · 60 s Free AI-readiness audit See your site the way ChatGPT does.
Personal reply within one business day.
CLIENT · NAME CHANGED

Law · Content engine + Edge Hub · 2026

How a German law firm put 1,100+ pages in front of AI search

Thousands of the firm’s articles sat on a third-party directory, building the directory’s authority. I built an engine that brings them home: rewritten for AI answers, checked claim by claim against the original, released through a one-click daily approval digest in pilot mode, published on the firm’s own domain.

Live · built Jul–Sep 2026 · handed over Sept 2026 · numbers observed 2026-10-05

01 · Before

The firm’s expertise was building someone else’s domain

The firm had done the hard part for years: writing expert answers to the exact questions its clients ask. But it published them on a directory profile. Every article made the directory stronger, and a buyer who asked ChatGPT or Perplexity was sent there, not to the firm. Copying thousands of articles by hand was out of the question. So was rebuilding a website the partners liked.

Before: who got the credit for the firm's expertise AI assistants and search engines cite a third-party legal directory, which holds thousands of the firm's articles. The firm's own domain has no knowledge section, so the link to the firm is broken. ASKED BY BUYERS AI assistants & search cites THIRD-PARTY Legal directory thousands of the firm’s articles live here THE FIRM Own domain no knowledge section to cite a contact form, nothing to quote
FIG 01 — Before: the expertise built someone else’s domain
Situation audit · beforestarting point
expertise
problem: Thousands of articles, published on a third-party directory
own domain
problem: No knowledge section, nothing for AI assistants to quote
AI crawlers
risk: Not measured. Nobody knew if GPTBot or ClaudeBot ever came
leads
risk: No way to tell which article brought a client in
website
note: WordPress, liked by the partners, to be left alone
risk
risk: Legal content: one invented claim is one too many
Brief: bring the expertise home, without touching the site and without a single unchecked claim.

Recognise this? See what AI sees on your site.Free AI-readiness audit · about 60 seconds · no email needed for the score.

02 · What I built

Two systems, one engine Gold is a human decision. Teal is an agent at work.

Part 1 · The content engine From a directory profile to approved pages on the firm’s domain, every five minutes

A Cloudflare Worker wakes every five minutes. It captures new and old articles from the directory, archives the raw page, and has Claude rewrite each one into a short, answer-first page with an FAQ. A second model then checks every factual claim against the original. Anything it can’t find in the source is flagged, and the draft is held for a person. Pages start in pilot mode: a lawyer releases each one with a click from the daily digest. The firm can switch categories it trusts to automatic publishing; any draft that fails the claim check always waits for a person.

  1. Step 1 (input): Directory profile, the firm’s articles.
  2. Step 2 (automated): Capture & archive, every 5 min.
  3. Step 3 (ai agent): Rewrite, answer-first + FAQ. Rejected here: Over budget (€100 breaker).
  4. Step 4 (ai agent): Claim check, vs. the source. Rejected here: Unsupported claims (flagged · held).
  5. Step 5 (human gate): Lawyer approves, pilot mode, one click.
  6. Step 6 (live): Published, firm’s own domain.
  7. Step 7 (automated): Crawler watch, 16 AI bots by name.
FIG 02 — The content engine, part 1 Seven steps, one human gate. The rewrite step stops dead at the €100 monthly budget breaker; the claim check flags anything the source does not support and holds that draft for a person.
PLATE 01 — Knowledge page with its claim-check rail, desktop · approval digest, phone · rebuilt mock, name changed

Part 2 · The Edge Hub One Word file became 230+ linked pages. WordPress was never touched.

A sister domain had its articles in a single Word document. I extracted them deterministically, let Claude write meta descriptions and check the structure, and served the pages from the edge inside the site’s live WordPress header and footer. Then an AI pass proposed internal links between the articles, and two gates threw out everything that wasn’t real or relevant.

  1. Step 1 (input): One Word file, 230+ articles.
  2. Step 2 (automated): Extract & render, deterministic.
  3. Step 3 (ai agent): Link suggestions, 1,175 proposed.
  4. Step 4 (automated): Validation gate, anchors + targets. Rejected here: −72 invented (anchors & pages).
  5. Step 5 (ai agent): Relevance check, second pass. Rejected here: −179 (irrelevant).
  6. Step 6 (human gate): Merge approved, reviewed change.
  7. Step 7 (live): Live in the WP shell, 924 links.
FIG 03 — The Edge Hub, part 2 AI proposes links; deterministic code checks that every anchor and every target exists; a second AI pass judges relevance; a human merges.

1,175 AI internal-link suggestions

Gate 1 · Deterministic validation −72 · 36 invented anchors · 36 links to pages that don’t exist
Gate 2 · Relevance check −179 · failed the relevance check
Gate 3 · Approved 924 published

What the AI proposed — and what survived

1,175 links proposed. 924 survived.

One Word file became 230+ articles. An AI model proposed internal links between them. Three gates checked every suggestion before anything went live — and every reject was logged, not hidden.

  1. 1,175 AI internal-link suggestions
  2. −72 hallucinated — 36 invented anchors, 36 links to pages that don’t exist
  3. −179 irrelevant — failed the relevance check
  4. 924 approved and published · avg 3.97 per article

Each dot is one real suggestion from the pipeline log.

Part 2 · Edge Hub internal links · Corvane Legal (client name changed)

03 · Who did what

Who did what The agents draft and check. A lawyer decides what goes live.

  1. agent://researcher

    Mapped the directory profile, the firm’s site and which AI crawlers could reach it.

    Never allowed to: decide what gets built or published.

  2. agent://architect

    Drafted the data model, the page template and a path-scoped route that adds the knowledge section without touching WordPress.

  3. human://vegardHUMAN GATE

    Approved the architecture and set the one rule that matters: no page goes live unless it passes the claim check, and nothing publishes automatically unless the firm has switched that category to auto mode.

  4. agent://writer

    Rewrites each article into an answer-first page with an FAQ, in the firm’s voice.

    Never allowed to: add a fact that is not in the source.

  5. agent://fact-checker

    Classifies every factual claim as supported, partial or unsupported against the original. Unsupported claims are flagged and the draft is held for a person.

    Never allowed to: let an unsupported claim through.

  6. human://clientHUMAN GATE

    A lawyer at the firm approves pages from the daily digest in pilot mode and decides which categories may publish automatically. One click per article; the link works once.

  7. agent://tester

    Runs 649 automated tests offline, against real captured pages, on every change.

  8. agent://watcher

    Counts 16 AI crawlers by name, flags when they stop visiting, and enforces the €100 monthly budget breaker.

    Never allowed to: spend past the cap.

  9. human://vegardHUMAN GATE

    Signed off every deploy, then moved the whole system into the firm’s own Cloudflare account in Sept 2026.

04 · Results

Results, with receipts Every number dated, with its method.

Law firm content engine: results with method and date
NumberWhat it measuresHow it was measuredObserved
1,100+✓ verified pages live on the firm's own domain live sitemap count 2026-10-05
230+✓ verified pages on a sister domain live sitemap count 2026-10-05
924✓ verified AI-suggested internal links that passed both gates and went live pipeline log 2026-10-05
72✓ verified hallucinated links rejected (36 invented anchors, 36 non-existent pages) deterministic validation log 2026-10-05
179✓ verified irrelevant links rejected by relevance check pipeline log 2026-10-05
649✓ verified automated tests test suite count 2026-10-05
16✓ verified AI crawlers tracked by name crawler module 2026-10-05
Not measured yet How often ChatGPT, Perplexity and Google’s AI answers cite the firm baseline being collected; first report at the 90-day refresh —
Not measured yet Enquiries attributed to a specific article attribution is live; too early to publish a count —

Page counts come from the live sitemaps. I don’t publish traffic or citation numbers until they have a full quarter behind them.

05 · Running costs

What it costs to run Capped, predictable, paid to the providers directly.

The CFO version: AI rewriting is throttled to 15 articles a day, which costs ~€33 a month. A hard €100 breaker stops every AI call if something goes wrong. Working through the entire back catalogue costs ~€230 in total, spread over roughly seven months. The bills go to the firm’s own accounts. I add no markup.

AI rewritingClaude, throttled to 15 rewrites a day
~€33/mo
Budget breakerhard monthly stop, set by the firm
€100
Back catalogue, one-offthe full rework, spread over about seven months
~€230
Scrapinga month in steady state, inside the provider’s 1,000-credit free tier
~330 credits
Source archiveevery captured page, kept for free re-extraction
~$0.13/mo

06 · Ownership

What the client owns

  • The Cloudflare accountWorkers, database, storage and routes moved into the firm’s account in Sept 2026.
  • The code and its CIFull repository; every push to main deploys from the firm’s own pipeline.
  • Every page and every source capturePublished pages, drafts, versions and the raw archive.
  • The AI keys and the budgetThe firm’s own Anthropic account. The €100 breaker is a setting it controls.
  • Tests and fixtures649 tests that run offline against captured pages.
  • Admin, runbook and rollbackAdmin behind Cloudflare Access, a written runbook, a rollback Worker kept in place.

07 · What I’d build for you

Turn the expertise you already wrote into pages AI can cite

If your firm has years of expert writing in Word files, PDFs, a directory profile or an old blog, I’ll build the same machine for you: capture, rework, fact-check, your approval, publish, measure. Start with an Edge Hub beside your current site, or go straight to the full engine.

  • Runs on your domain and in your Cloudflare account from day one
  • Every claim checked against your source; you approve pages in pilot mode and choose which categories go automatic
  • A hard monthly AI budget that you set
  • Per-bot AI-crawler analytics and article-to-enquiry attribution

Not for you if you want volume without review. Every page passes the claim check, and a person decides what publishes by hand and what may publish automatically.

Mapped offer · Cited

Content Engine · Full Engine

from €9,500 · $10,500 typical €9.5k–18k

6–9 weeks · Continuous capture → rework → fact-check → approval → publish.

Or start smaller · Cited

Content Engine · Edge Hub

from €4,500 · $4,950 typical €4.5k–7k

2–3 weeks · Your back catalogue as a knowledge hub beside your current site.

Not sure which fits? Book 20 minutes with me
08 · Under the hood The technical depth, for your CTO
Runtime
Cloudflare Workers in TypeScript, no framework. D1 with 36 migrations, KV for pages and the raw archive, Cron every 5 min with a scheduler lock and a heartbeat ping to an external monitor.
Pipeline
scan → fetch → parse → rework → verify → publish → digest. The parser fails loudly: if the source markup changes, the source is flagged and an alert goes out. Nothing is guessed.
Models
Claude Sonnet rewrites; Claude Haiku writes meta descriptions and runs structural QA. JSON-LD is assembled in code, never by a model.
Approvals
Approval links are random, single-use, bound to one article, stored only as a hash, and expire after 72 h. Client preview links are revocable.
Edge overlay
The Worker is bound to one path route, so it cannot touch any other URL. If it throws, the request passes through to WordPress. The Edge Hub borrows the live header, footer and cookie banner and caches that shell for 1 h.
Link validation
Every AI-suggested link must use an anchor that exists word for word in the article and point to a page that exists. 72 failed: 36 invented anchors, 36 non-existent pages. A relevance pass then rejected 179 more, leaving an average of 3.97 links per article.
AI crawlers
Explicit allowlist, llms.txt, sitemaps, IndexNow, and per-bot hit counts for 16 named AI crawlers.
Attribution
The article a visitor read travels in the URL to the contact form. No cookies, no storage.
Tests
649 tests in 42 files, runnable offline in mock mode against real captured pages.
Handover
Zero-downtime move into the firm’s own Cloudflare account in Sept 2026, CI deploys on push, rollback Worker kept in place.