← All projects
Gated · private

Voice Engine

A measured style guide built from evidence instead of impression — two corpora, 330,000 words, a statistical control group, and three failures the process caught in itself.

330K+ wordsmeasured, not inferred
2 pipelinesbuilt independently
0 serversfiles + scripts
Data flow

Two pipelines built months apart that converge only at the intersection — plus the feedback edge where the clean speech control corrected the mail corpus.

The interactive data-flow diagram is built for a larger screen.
Open the full diagram in its own tab ↗

Why it exists

Every AI tool will happily write "in your voice", and none of them know what that voice is — they infer it from whatever has been typed into the box. Ask an LLM to describe how you write and it produces something plausible and unfalsifiable. The purpose here was to measure it instead: build a style guide from evidence, where every rule carries a ratio against a control group or does not ship, and the output is prompt blocks that make any LLM draft like a specific person.

What it is

Two independent corpora — seven years of sent mail pulled through the Microsoft Graph API, and roughly 100 long-form recorded conversations — built months apart with no shared measurement code, because agreement across two pipelines is evidence and agreement within one is a bug you have not found yet. The mail corpus is cleaned and collapsed with MinHash near-duplicate detection, which revealed that 57% of it was templated and reordered the whole dataset. Registers and eras are derived from observable features rather than assumed. The speech corpus carries something the mail corpus could not: a near-perfect control group in the counterpart turns — same topic, same room, minutes apart — which turns every lexical claim into a measured gap against how peers talk in identical conditions. That control corrected the mail corpus after the fact, which is the most important arrow in the system. Everything is scored through a blind validation harness that makes withheld metadata structurally unavailable rather than merely off-limits. What survives both corpora ships as three standalone prompt blocks plus machine-readable rule sets. The one-line finding: functions transfer, phrasings do not.

The hard part

The brag is not the output. It is that the process caught itself being wrong three times, and each catch generalises to measuring anything from a corpus you are too close to.

The subject could not pass his own validation harness. Blind drafts scored 2.93 out of 8 against a 5.6 pass bar. Before reporting that as a failure, the build calibrated the harness on known-good input — and found that a real turn by the actual person scored 3.05. Ceiling-to-floor separation was 0.76 points out of 8; the harness was measuring almost nothing. Two defects, both in the spec: a move-type overlap metric with a median of zero between two genuine turns by the same person, and content-word thresholds imported from a different medium that sat above the real ceiling. Rebuilt with every threshold re-derived from measured distributions, the ceiling went to 4.48 and blind drafts to 4.20 — 94% of ceiling. Still reported as a number rather than a pass, because a 5.6 bar cannot be applied to a harness whose ceiling is 4.48. Calibrate the instrument on ground truth before you trust any score it produces.

Unconditional rates are backwards for anything worth drafting. The guide reported a 26.5% greeting rate — correct corpus-wide, and useless, because the median message is 34 words and the largest register is a 16-word bare reply. For messages over 40 words, which is every message anyone would ask an assistant to write, the real greeting rate is 75%. Openings and sign-offs accounted for 17 of 26 zero-scores in validation. Condition every rate on the situation the rule will be applied in.

The canonical document contradicted its own summary. The portable prompt block still claimed one-sentence paragraphs from a corpus-wide median of 1, while a later section in the same file refuted that as a pooling artefact and gave the real per-message figure. Both shipped artifacts disagreed on the one mechanic that was failing validation. Grep for the claim, not just the section.

Underneath all three sits the design constraint that made them findable: every rule carries a ratio against a control, or it does not ship. The second corpus exposed a systematic defect in the first — a contaminated baseline does not understate distinctiveness, it inverts the ranking. Words the first guide called characteristic turned out to be used at the same rate by ~100 different professionals in identical conditions; words it ranked lowest turned out to be the real markers. And the largest finding was not lexical at all but grammatical, reliable enough to be used forensically when one transcript arrived with scrambled speaker labels.

Worth highlighting

Functions transfer, phrasings do not

Only rules independently attested in both corpora survive into the default prompt block — eight conversational moves and eight words. What ships is three standalone prompt blocks plus machine-readable rule sets, each written to work pasted into a tool that has never heard of the project.

A harness that can't peek

Blind drafts are scored by a validation harness that emits only the metadata a drafter is allowed to see and seals the rest mechanically. Withheld information is structurally unavailable rather than merely off-limits, and every threshold is derived from measured distributions rather than chosen.

vibe-coded with ClaudeA measured style guide, vibe-coded end to end — and the process caught itself being wrong three times before anything shipped.

More projects

← Back to all projects