We gave two frontier AI models the same brief. Here is what they built.

by Ross Gordon, Founder, Assist IQ

Last night I ran an experiment I could not have run six months ago.

I gave two frontier AI models the exact same brief. Same design system. Same page-by-page spec. Same instruction: build a complete website, optimised for search and for the AI answer engines people now use instead of Google. Then I let them build.

One was Anthropic's Fable 5. The other was OpenAI's GPT-5.6, codename SOL. Not a chat. Not "help me with a section." A full 14-page site each, plus the legal pages, built end to end, in a single shot.

What both models got right

Both produced a complete site. AI employees, managed operations, service pages, industry pages, proof, an honest competitor comparison, a pillar guide, about, FAQ, contact, plus the legal pages. Real copy, not filler. Valid structured data on every page. Answer-first content blocks written for how people now search. Internal linking across the whole thing. And both compiled cleanly on the first build, with zero errors.

Six months ago that was science fiction. You would get something that looked plausible, then watched it fall over the moment you tried to build it. Broken links, invented client names, inconsistent design, markup that did not validate. You would lose a day fixing it. Last night I spent the evening choosing between two sites that were both 90 percent of the way there.

Where they differed

Fable's strength was design discipline. It used the full range of the system: image blocks, cards, comparison tables, alternating layouts. It read like an agency built it. It also did something I did not expect. It found my feedback from an earlier round sitting in the project files and applied it without being asked. That is judgment, not just output.

GPT-5.6 SOL was the surprise. Genuinely close. Sharp, clean copy, arguably tighter in a few places. And relentlessly rigorous. It ran its own quality sweep and reported zero broken links, zero invalid markup, zero of the formatting I had banned. Its weakness was that the inner pages leaned text heavy, and it did not pick up the unwritten rules the way Fable did. It needed one cleanup pass.

So Fable felt like the designer. SOL felt like the meticulous engineer.

The verdict

My verdict, and it is only my opinion: Fable is the stronger of the two. But SOL is not far behind. Honestly the gap was small enough that I kept changing my mind clicking through them. Both feel like a clear generation above Opus 4.8, the model that wrote the brief and ran the whole comparison in the first place.

Here is the part that stays with me. If this is the worst these models will ever be, and they only get better from here, then the one-shot production website is not coming. It is already here. Not push-button yet. It still needed a human for the calls that matter: which logo, which testimonial, the final sign-off. But 90 percent, in an evening, from a brief.

See for yourself

Both original one-shot builds are live to click through:

The site you are reading this on is the shipped result: Fable's build, with a few of Sol's best bits grafted in, and the human calls made by a person.

That is the whole point of how we work. The machine carries the load. Judgment stays with someone accountable. If you want that applied to your own business, the first step is a straight conversation about what is actually slowing you down. Start with the AI Operations Day.

Same brief. Same spec. Two models. One evening. Which one would you have shipped?

More from the blog

Claude Code for business: what a real deployment involves

Deploying Claude Code across a company is eight decisions, not an install. Here is the deployment decision document we use, the security questions a client will ask you, and what Anthropic assesses before it issues its Claude Code partner badge.

Read more

Claude certification: what Anthropic actually tests

The Claude Certified Associate exam is proctored, closed book and scored out of 1000. The surprise is what it weights most heavily: not writing prompts, but knowing when the answer is wrong and when a human has to check it.

Read more

Start with a free enquiry

Not a discovery call. Not a pitch with a calendar link. Five questions about how your business actually runs, and it costs nothing. Ross reads every enquiry himself and replies within one working day with a straight first answer: what looks worth automating, and what doesn’t.

If it looks like we can genuinely help, the next step is the AI Operations Day: one working day inside the business, followed by a written Opportunity Map. It shows what is hurting, what should stay human, and the best one or two jobs to prove first. That part comes later, and only if it makes sense for you.

Tell us what is slowing you down

Free to ask. No obligation. Ross replies personally within one working day.