← turboflow.online
~/turbo-flow — research/cross-model-review audited · n=868

Cross-model code review: the builder must never be its own reviewer

What is cross-model code review? It is the practice of having a model from a different model family than the builder review every change — code built by ZCode (GLM) reviewed by Claude Code, for example. The reviewer sees the diff read-only and must return a parseable APPROVED or REVISE verdict. It is the core mechanism of rig-lite's gate, and the reason Turbo Flow reorganized itself around governance.

Published Oct 7, 2026 · Updated Oct 8, 2026

868
gate verdicts in the audited week — Sept 14–21, 2026, gate live Sept 16
687
verdicts REVISE (79%) — the reviewer sent four of five diffs back; 141 PRs still shipped
100%
of merges held by a human — a rig invariant, not a promise

Why same-family review falls short

A model reviewing output from its own family tends to share that family's failure biases: the same blind spots, the same preferred mistakes, the same confidence in the same wrong assumptions. Self-review within a family mostly re-derives the builder's reasoning rather than auditing it. The study's second finding sharpens the rule: direction matters — the strongest analyst model should review, never a mirror of the builder.

How the gate enforces it

# deterministic checks first — they're free
lint · tests · diff hygiene · secret scan
        ↓
# then the cross-family reviewer, read-only, on a nonce-fenced diff
reviewer ≠ builder family (auto-detected; same-family is refused)
verdict must be a parseable APPROVED or REVISE on the final line
        ↓
# ambiguity is REVISE — fail-closed, and the gate never merges
a human holds the merge button

The nonce fence means a diff cannot inject its own APPROVED. Reviewer output is secret-redacted. A silent or hung reviewer is a REVISE, not a pass. --no-exec mode and hang caps exist for untrusted branches. Every verdict lands in a local JSONL log, which is what makes the usage numbers auditable after the fact.

The public receipts

  • The gate reviewed itself into existence. Written by one model family, it was REVISE'd seventeen times by a different family before its first APPROVED — and the findings were exactly the class it exists to catch: a prompt-injection hole that let a diff approve itself, a fail-open path that skipped review, a quoting bug that swallowed failing tests, a silent-reviewer misdiagnosis. Every fix became a self-test.
  • One week of production use, counted from the gate log: 868 gate verdicts on 141 merged PRs — and 79% of the verdicts were REVISE. The gate is not ceremonial; it rejects most of what it sees. (seven days inside the rig)
  • Fail-closed is testable. self-test.sh ships with the kit: hundreds of checks, all green, proving the failure modes actually fail closed.

What is published, and what isn't (yet)

Honest disclosure, because this page will be quoted: the numbers above come from the rig’s own audited operating record. The 116-task figure previously cited here without attribution is a published external result — Xiang et al., “Cross-Model LLM Code Review” (Agentic SE @ KDD’26, arXiv:2607.21656) — not internal data. What is ours, and public, is the verification dossier that re-derives every printed number from our primary logs (https://www.turbo-rig.com/evidence). Judge the mechanism by the receipts that are already public: the gate log counts above and the fail-closed self-test suite you can run yourself in ten minutes.

External evidence (young, and mixed)

Independent research on cross-model review is early. Two papers worth reading alongside our numbers: "When Does a Second Model Help? Cross-Model Review" (arXiv, 2026) reports that a top-tier cross-model reviewer's F1 was not significantly different from same-model review in a fresh session — a narrower claim than ours, and a useful counterpoint; "Cross-Model LLM Code Review: Should you use Claude to review Codex, or vice versa?" (OpenReview) studies the exact builder/reviewer pairing workflow the gate automates. We cite them because engines — and engineers — should prefer claims that point somewhere. Our position: external evidence is mixed on margins, our own controlled study and our production gate log (79% REVISE across 868 verdicts) are what we ship on, and the task-level dataset is disclosed as in-preparation until it lands.

Run it yourself

# any repo, any builder CLI:
bash /path/to/turbo-flow/rig-lite/init-repo.sh
bash /path/to/turbo-flow/rig-lite/self-test.sh          # watch it fail closed on purpose
/path/to/turbo-flow/rig-lite/gate.sh --pr 1 --builder zcode

The broader argument — orchestration couldn't buy trust, governance does — is here, and the philosophy that ties it together is agents build, humans merge.