← turboflow.online
~/turbo-flow — research/ notes · data · experiments

The evidence base

This project makes claims that can be checked, so the checking is part of the project. Research here means three things: measured production usage (the seven-day window below, straight off the gate's own JSONL log), a public audit of every printed number (the verification dossier at turbo-rig.com/evidence), and the practitioner findings the architecture stands on (orchestration vs governance).

Published Oct 7, 2026 · Updated Oct 8, 2026

141
PRs merged in the seven-day window · across 5 repos · every merge human
868
gate verdicts · gate live Sept 16–21
79%
of those verdicts were REVISE
104
parallel worktree lanes
4.21B
tokens through the engine
$416
total review spend for the week

Seven days inside an agentic coding rig

One week — September 14–21, 2026 — building with the rig that builds itself. The numbers above are not a benchmark someone ran; they are the week's operating record, and the full breakdown lives at turbo-rig-stats-only-sept14-21.vercel.app. Three of them deserve interpretation:

  • 79% of gate verdicts were REVISE. This is the single most important number on this page. A review gate that approved everything would be ceremony; a gate that rejects roughly four of every five diffs it sees is doing the actual work of quality control — catching prompt-injection holes, fail-open paths, silent test failures, overclaims — before a human ever looks. It is the strongest available evidence that "agents build, humans merge" is an operating system and not a slogan.
  • 868 verdicts over 141 merges ≈ 6 gate rounds per merged PR (verdicts cover the gate-live window, merges the full week — mixed basis, stated). Changes go through the gate repeatedly: build → REVISE → fix → REVISE → fix → APPROVED → human merges. The revision loop is where the quality lives; the 79% REVISE rate is what powers it.
  • 104 worktree lanes. Parallel writers lived in isolated worktrees and converged through the gate — the isolation setup that the merge-tax numbers (336k conflicts across 142k agent PRs elsewhere) say is the prerequisite for parallel agents at all.

The spend side: 4.21B tokens and $416 of review cost for a week that shipped 141 production merges across five repos — governance is not free, but at roughly $3 of review per merged PR (spend covers the gate-live window, merges the full week) it is dramatically cheaper than what it replaces.

The papers & notes

Cross-model code review — the audited record

Builder family ≠ reviewer family. In the rig’s first audited week: 687 of 868 gate verdicts REVISE (79%) while 141 PRs still shipped. Every number re-derived in the public verification dossier.

read the evidence page →

Agent orchestration vs agent governance

Why coordination doesn't scale without trust: the AgenticFlict numbers, the harness absorption wave, and "the infrastructure is the coordination".

read the comparison →

The gate's own review record

The review gate was itself REVISE'd cross-family seventeen times before earning APPROVED — and the findings (prompt injection, fail-open paths, quoting bugs) each became a permanent self-test. The record is on the evidence page.

see the receipts →

Methodology disclosure

Where each number comes from: the seven-day figures are counted from the gate's local JSONL verdict log joined against the GitHub merge record, and mirrored on the public stats page linked above and on turbo-rig.com/evidence. Every claim was re-derived in an independent audit whose dossier — tables, corrections, reproduction commands — is public there. The 116-task figure previously cited here without attribution is a published external result (Xiang et al., Agentic SE @ KDD’26, arXiv:2607.21656), not internal data — cited as such from now on. Usage numbers on this site are refreshed when a window is published, not sampled live — a number on a Turbo Flow page always traces to the artifact that wrote it.