We built 215 tools and 60 agents.
Then we deleted almost all of it.
Turbo Flow v4 (March 2026) was, by tool-count standards, a serious agentic development environment: 215+ MCP tools, 60+ agents, swarm orchestration via Ruflo, cross-session memory, a codebase knowledge graph, per-agent worktree isolation. Seven months later, v5.2 shipped 35 files — and the orchestration stack was gone. This page is the honest story of why. It is the same story the version history tells and the usage data backs up.
Published Oct 7, 2026 · Updated Oct 8, 2026
The original hypothesis
More specialized agents + more tools + orchestration = more autonomous engineering. It was not a silly idea — it was the consensus of the field in 2025–2026, and the v4 numbers were real:
- 215+ MCP tools wired into every session (175+ before the Ruflo migration)
- 60+ agents, from coder to architect to reviewer roles
- Ruflo 3-tier model auto-routing — roughly a 75% cut in model spend
- Beads memory so agents survived session boundaries; GitNexus so they knew the codebase
What actually happened
- Complexity grew faster than capability. Every new tool and agent was another integration to install, wire, verify, and keep current. The setup script alone became a project.
- Coordination became expensive. Multi-agent output still had to converge on one branch, and the merge tax grew with the swarm. (This is not just our anecdote: analysis of 142k agent PRs — AgenticFlict — counted 336k merge conflicts.)
- Humans couldn't meaningfully review the volume. Faster agent output with a human as the only check just moves the bottleneck to the least scalable part of the system.
- The harnesses got smarter. Swarms, memory, routing, subagents — the things the orchestration stack existed to provide — started shipping natively in the coding harnesses themselves. The wrapper had lost its reason to exist.
The realization
The missing layer wasn't more intelligence. It was governance.
Once the models are competent and the harness runs the loop, the question that decides whether you can actually ship is: who verifies the work? A model reviewing its own family's output shares that family's blind spots — so the reviewer always comes from another family: a design choice, not something that week measured. What it measured is the gate's strictness — 687 of 868 diffs sent back (gate live Sept 16–21) while 141 PRs still shipped (the record is on the cross-model code review page). That design choice, plus a season of feeling the merge tax, rewrote the roadmap.
Therefore — the pivot
Ruflo / swarm orchestration ← deleted (harnesses do this natively now)
215+ tools, 60+ agents ← deleted (same reason; recoverable in v1.0.1 → v4.0 tags)
↓ replaced by
constitution.md ← the laws every agent operates under
gate.sh (cross-model review) ← builder ≠ reviewer, fail-closed, never merges
memory/ (git-versioned) ← agents forget; git doesn't
wt.sh (worktree isolation) ← parallel writers, no merge collisions
human merge authority ← the one thing that never moved
↓ becoming
TURBO RIG ← the integrated system, private beta
What survived, and where this goes
Turbo Flow lives on as the open-source rules layer: rig-lite, a governance kit that works with any harness — because it doesn't try to be the harness. Turbo Rig is the same ideas integrated into one always-on system; the comparison page draws the line precisely. The philosophy underneath both is one sentence: agents build, humans merge.