Play: I Put Fable 5.1 Up Against GPT6-Astra And More
Devlog

I Put Fable 5.1 Up Against GPT6-Astra And More

9:11 • Published September 5, 2026 • Watch on YouTube ↗

Fable 5.1, GPT 5.6, DeepSeek V4 Pro, GLM 5.3, and GPT 6: I gave all five the same brief to build an agentic IDE for Ollama. Then I tried to actually use what they made.

The project is called Yarn. Each model got room to invent its own workflow, with minimal extra context and follow-up rounds to fix problems. Some results were polished. Others opened but barely worked. One took about 14 hours of elapsed time, including waits for usage limits to reset.

I compare the apps, the fixes they needed, and whether the results justified the time and model usage. The model that built my favorite app isn’t necessarily the one I’d choose for this job again.

Explore the Yarn project and compare all five builds.

Browse the complete Yarn source code.

All five implementations are open source under the MIT license. Try them, compare the code, or build on them.

Chapters 0:00 Why are there four versions of Yarn? 0:48 The prompt and setup 1:18 DeepSeek V4 Pro: fast, with functional gaps 2:03 GLM 5.3: an app that wouldn’t work 2:54 GPT 5.6: the efficiency standout 4:33 Fable 5.1: capable, but worth the wait? 7:13 What I’d use for everyday coding 7:43 GPT 6 enters the comparison 8:29 My final pick 8:55 Get the code

These are my results from one app-building experiment, using Claude Desktop, Codex, and OpenCode, including follow-up fixes.

Which model would you trust to build this?

← Back to all videos