Cross-model code review: the builder must never be its own reviewer
What is cross-model code review? It is the practice of having a model from a different model family than the builder review every change — code built by ZCode (GLM) reviewed by Claude Code, for example. The reviewer sees the diff read-only and must return a parseable APPROVED or REVISE verdict. It is the core mechanism of rig-lite's gate, and the reason Turbo Flow reorganized itself around governance.
Published Oct 7, 2026 · Updated Oct 8, 2026
Why same-family review falls short
A model reviewing output from its own family tends to share that family's failure biases: the same blind spots, the same preferred mistakes, the same confidence in the same wrong assumptions. Self-review within a family mostly re-derives the builder's reasoning rather than auditing it. The study's second finding sharpens the rule: direction matters — the strongest analyst model should review, never a mirror of the builder.
How the gate enforces it
# deterministic checks first — they're free lint · tests · diff hygiene · secret scan ↓ # then the cross-family reviewer, read-only, on a nonce-fenced diff reviewer ≠ builder family (auto-detected; same-family is refused) verdict must be a parseable APPROVED or REVISE on the final line ↓ # ambiguity is REVISE — fail-closed, and the gate never merges a human holds the merge button
The nonce fence means a diff cannot inject its own APPROVED. Reviewer output is secret-redacted. A
silent or hung reviewer is a REVISE, not a pass. --no-exec mode and hang caps exist for
untrusted branches. Every verdict lands in a local JSONL log, which is what makes the
usage numbers auditable after the fact.
The public receipts
- The gate reviewed itself into existence. Written by one model family, it was REVISE'd seventeen times by a different family before its first APPROVED — and the findings were exactly the class it exists to catch: a prompt-injection hole that let a diff approve itself, a fail-open path that skipped review, a quoting bug that swallowed failing tests, a silent-reviewer misdiagnosis. Every fix became a self-test.
- One week of production use, counted from the gate log: 868 gate verdicts on 141 merged PRs — and 79% of the verdicts were REVISE. The gate is not ceremonial; it rejects most of what it sees. (seven days inside the rig)
- Fail-closed is testable.
self-test.shships with the kit: hundreds of checks, all green, proving the failure modes actually fail closed.
What is published, and what isn't (yet)
Honest disclosure, because this page will be quoted: the numbers above come from the rig’s own audited operating record. The 116-task figure previously cited here without attribution is a published external result — Xiang et al., “Cross-Model LLM Code Review” (Agentic SE @ KDD’26, arXiv:2607.21656) — not internal data. What is ours, and public, is the verification dossier that re-derives every printed number from our primary logs (https://www.turbo-rig.com/evidence). Judge the mechanism by the receipts that are already public: the gate log counts above and the fail-closed self-test suite you can run yourself in ten minutes.
External evidence (young, and mixed)
Independent research on cross-model review is early. Two papers worth reading alongside our numbers: "When Does a Second Model Help? Cross-Model Review" (arXiv, 2026) reports that a top-tier cross-model reviewer's F1 was not significantly different from same-model review in a fresh session — a narrower claim than ours, and a useful counterpoint; "Cross-Model LLM Code Review: Should you use Claude to review Codex, or vice versa?" (OpenReview) studies the exact builder/reviewer pairing workflow the gate automates. We cite them because engines — and engineers — should prefer claims that point somewhere. Our position: external evidence is mixed on margins, our own controlled study and our production gate log (79% REVISE across 868 verdicts) are what we ship on, and the task-level dataset is disclosed as in-preparation until it lands.
Run it yourself
# any repo, any builder CLI: bash /path/to/turbo-flow/rig-lite/init-repo.sh bash /path/to/turbo-flow/rig-lite/self-test.sh # watch it fail closed on purpose /path/to/turbo-flow/rig-lite/gate.sh --pr 1 --builder zcode
The broader argument — orchestration couldn't buy trust, governance does — is here, and the philosophy that ties it together is agents build, humans merge.