We gave two frontier AI models the same brief. Here is what they built.

by Ross Gordon, Founder, Assist IQ

Last night I ran an experiment I could not have run six months ago.

I gave two frontier AI models the exact same brief. Same design system. Same page-by-page spec. Same instruction: build a complete website, optimised for search and for the AI answer engines people now use instead of Google. Then I let them build.

One was Anthropic's Fable 5. The other was OpenAI's GPT-5.6, codename SOL. Not a chat. Not "help me with a section." A full 14-page site each, plus the legal pages, built end to end, in a single shot.

What both models got right

Both produced a complete site. AI employees, managed operations, service pages, industry pages, proof, an honest competitor comparison, a pillar guide, about, FAQ, contact, plus the legal pages. Real copy, not filler. Valid structured data on every page. Answer-first content blocks written for how people now search. Internal linking across the whole thing. And both compiled cleanly on the first build, with zero errors.

Six months ago that was science fiction. You would get something that looked plausible, then watched it fall over the moment you tried to build it. Broken links, invented client names, inconsistent design, markup that did not validate. You would lose a day fixing it. Last night I spent the evening choosing between two sites that were both 90 percent of the way there.

Where they differed

Fable's strength was design discipline. It used the full range of the system: image blocks, cards, comparison tables, alternating layouts. It read like an agency built it. It also did something I did not expect. It found my feedback from an earlier round sitting in the project files and applied it without being asked. That is judgment, not just output.

GPT-5.6 SOL was the surprise. Genuinely close. Sharp, clean copy, arguably tighter in a few places. And relentlessly rigorous. It ran its own quality sweep and reported zero broken links, zero invalid markup, zero of the formatting I had banned. Its weakness was that the inner pages leaned text heavy, and it did not pick up the unwritten rules the way Fable did. It needed one cleanup pass.

So Fable felt like the designer. SOL felt like the meticulous engineer.

The verdict

My verdict, and it is only my opinion: Fable is the stronger of the two. But SOL is not far behind. Honestly the gap was small enough that I kept changing my mind clicking through them. Both feel like a clear generation above Opus 4.8, the model that wrote the brief and ran the whole comparison in the first place.

Here is the part that stays with me. If this is the worst these models will ever be, and they only get better from here, then the one-shot production website is not coming. It is already here. Not push-button yet. It still needed a human for the calls that matter: which logo, which testimonial, the final sign-off. But 90 percent, in an evening, from a brief.

See for yourself

Both original one-shot builds are live to click through:

The site you are reading this on is the shipped result: Fable's build, with a few of Sol's best bits grafted in, and the human calls made by a person.

That is the whole point of how we work. The machine carries the load. Judgment stays with someone accountable. If you want that applied to your own business, the first step is a straight conversation about what is actually slowing you down. Start with the AI Operations Day.

Same brief. Same spec. Two models. One evening. Which one would you have shipped?

More from the blog

We ran an AI visibility report on our own brand: cited the most, recommended less

On 10 September 2026 we ran an AI visibility report on our own will-writing business: 20 prompts, four assistants, 80 measured cells. The domain was the most-cited domain in the sample and the brand was named far less often.

Read more

We moved our AI operator to Claude Fable 5.1 the day after release and tracked it for a week

Anthropic released Claude Fable 5.1 on 1 September 2026. We switched our own AI operator onto it the next day and tracked seven days of usage: 160.8M tokens a day against an 89.3M baseline, and nothing throttled all week, on a counter we have never yet seen fire.

Read more

Start with a free enquiry

Not a discovery call. Not a pitch with a calendar link. Five questions about how your business actually runs, and it costs nothing. Ross reads every enquiry himself and replies within one working day with a straight first answer: what looks worth automating, and what doesn’t.

If it looks like we can genuinely help, the next step is the AI Operations Day: one working day inside the business, followed by a written Opportunity Map. It shows what is hurting, what should stay human, and the best one or two jobs to prove first. That part comes later, and only if it makes sense for you.

Tell us what is slowing you down

Free to ask. No obligation. Ross replies personally within one working day.