← turboflow.online
~/turbo-flow — docs/best-practices from the gate log

AI code review best practices — from a gate that rejects 79%

In one sentence: put deterministic checks first, make the model reviewer come from a different family than the builder, demand a parseable verdict, and fail closed with a human on merge. Everything below comes from running exactly that in production — the cross-model gate whose log recorded 868 verdicts in one week, 79% of them REVISE — plus the measured window behind it.

Published Oct 9, 2026 · Updated Oct 9, 2026

1. Deterministic before expensive

Lint, tests, type-check, syntax, secret scan — these are free and deterministic, and a diff that fails tsc does not need a model to say so. Running them first is both an economics rule and a signal rule: when the paid reviewer speaks, it speaks about things only it can judge.

2. Reviewer family ≠ builder family

A model reviewing its own family's output shares that family's blind spots. The controlled evidence: cross-family review lifted pass rates from 71.6% to 89.7% (n=116) while same-family self-review barely moved. Direction matters too — the strongest analyst model reviews; never a mirror of the builder.

3. Verdicts are parseable or they don't exist

The reviewer must end with exactly VERDICT: APPROVED or VERDICT: REVISE. Ambiguity, hedging, a missing line — all REVISE. This one convention turns a chatty model into a component a script can act on, and it is why the gate can log 868 verdicts you can actually count.

4. Review small diffs

The gate's own history is the argument: its early REVISEs included a quoting bug that swallowed failing tests and a fail-open path hidden in volume. Small diffs get real reviews; 2,000-line walls get skimmed by machines and humans alike. Keep PRs small enough to be rejectable.

5. Fail closed, always

Silent reviewer = REVISE. Timeout = REVISE. Unparseable output = REVISE. A review gate that fails open is a rubber stamp with downtime. Every failure mode should default to "not approved," and the merge button stays with a human — agents build, humans merge.

6. Log every verdict

Every verdict lands in a local JSONL log. That log is what turned one week of building into evidence (868 verdicts, 79% REVISE, ~6 rounds per shipped PR) and what lets the digest join review state with the merge queue. If your review process leaves no artifact, you cannot learn from it.

Run this today

These six practices ship as rig-lite — bash-only, any repo, ten minutes, with a fail-closed self-test suite that proves each rule holds. For the full evidence base behind rules 2 and 5, see the cross-model review study page and the seven-day window.