We gave two frontier AI models the same brief. Here is what they built.
by Ross Gordon, Founder, Assist IQ
Last night I ran an experiment I could not have run six months ago.
I gave two frontier AI models the exact same brief. Same design system. Same page-by-page spec. Same instruction: build a complete website, optimised for search and for the AI answer engines people now use instead of Google. Then I let them build.
One was Anthropic's Fable 5. The other was OpenAI's GPT-5.6, codename SOL. Not a chat. Not "help me with a section." A full 14-page site each, plus the legal pages, built end to end, in a single shot.
What both models got right
Both produced a complete site. AI employees, managed operations, service pages, industry pages, proof, an honest competitor comparison, a pillar guide, about, FAQ, contact, plus the legal pages. Real copy, not filler. Valid structured data on every page. Answer-first content blocks written for how people now search. Internal linking across the whole thing. And both compiled cleanly on the first build, with zero errors.
Six months ago that was science fiction. You would get something that looked plausible, then watched it fall over the moment you tried to build it. Broken links, invented client names, inconsistent design, markup that did not validate. You would lose a day fixing it. Last night I spent the evening choosing between two sites that were both 90 percent of the way there.
Where they differed
Fable's strength was design discipline. It used the full range of the system: image blocks, cards, comparison tables, alternating layouts. It read like an agency built it. It also did something I did not expect. It found my feedback from an earlier round sitting in the project files and applied it without being asked. That is judgment, not just output.
GPT-5.6 SOL was the surprise. Genuinely close. Sharp, clean copy, arguably tighter in a few places. And relentlessly rigorous. It ran its own quality sweep and reported zero broken links, zero invalid markup, zero of the formatting I had banned. Its weakness was that the inner pages leaned text heavy, and it did not pick up the unwritten rules the way Fable did. It needed one cleanup pass.
So Fable felt like the designer. SOL felt like the meticulous engineer.
The verdict
My verdict, and it is only my opinion: Fable is the stronger of the two. But SOL is not far behind. Honestly the gap was small enough that I kept changing my mind clicking through them. Both feel like a clear generation above Opus 4.8, the model that wrote the brief and ran the whole comparison in the first place.
Here is the part that stays with me. If this is the worst these models will ever be, and they only get better from here, then the one-shot production website is not coming. It is already here. Not push-button yet. It still needed a human for the calls that matter: which logo, which testimonial, the final sign-off. But 90 percent, in an evening, from a brief.
See for yourself
Both original one-shot builds are live to click through:
- Fable: proposals.assistiq.co.uk/aiq-fable-oneshot
- GPT-5.6 SOL: proposals.assistiq.co.uk/aiq-sol-oneshot
The site you are reading this on is the shipped result: Fable's build, with a few of Sol's best bits grafted in, and the human calls made by a person.
That is the whole point of how we work. The machine carries the load. Judgment stays with someone accountable. If you want that applied to your own business, the first step is a straight conversation about what is actually slowing you down. Start with the AI Operations Day.
Same brief. Same spec. Two models. One evening. Which one would you have shipped?