My system redesign forked three times. Instead of crowning a favourite, I brought in the newest Claude model — cold, no history — and made it score all three against real data and choose the one that deserved to live.
The AIOS redesign didn't fork once — it forked three times. 1ultraplan, 2ultrathink, 3merged: three parallel prototypes, each a full app racing to replace the same running system. No scorecard, no clear winner, and drift creeping in as each one grew its own ideas.
The builder is the worst judge of his own forks — too much sunk cost in each. So I brought in Fable 5, the newest Claude model, completely cold. In one session it scouted all three prototypes, the production AIOS, and the architecture doc, then scored them on the same axes: real data, feature coverage, and drift.
3merged won outright. It had already absorbed 2ultrathink's entire app — capture spine, classifier, FTS5 search — plus 1ultraplan's real-data importer and the design-system skin, and it was the only prototype running on real data. But the fresh model flagged something a tired builder would have missed.
The obvious next move was to rebuild — a fourth, cleaner app. The fresh model said no. The real leverage is the Core: one SQLite spine with import/export. UI comes second. The AI that built all three versions of itself told its builder to stop over-building.
The diagram, the logged decision, the two archived prototypes, and the git graph across all three repos. Tap any image to enlarge it and read the exact prompt that drew it.




Decision logged with full reasoning, alternatives considered, and a rollback plan. The migration is parked to August — July stays on shipping value for the companies I work with. The Core gets built first; the UI second. The point of the exercise wasn't to crown a winner; it was to stop building the wrong thing.
The full story — three forked prototypes, a cold judge model, and the verdict that said stop building.
When you can't pick between your own prototypes, bring in a model that's never seen them and make it score against real data — not vibes. Everything in this episode is free and open — clone it, run it, make it yours.
# fresh-judge + decision-log pattern — repo publishing soon
Repo's being cleaned up for release. Comment JUDGE on the post and the bot DMs you the link the moment it's public.
One SQLite spine, import and export — the thing the judge said actually mattered. In August the winner stops being a plan and starts being the system.