Loading...
Loading...
We gave four leading AI models the same 50 small-business briefs and asked each one for a Facebook ad, twice. 400 ads came back. The models followed instructions almost perfectly. That turned out to be the problem.
By Ad Legends ·
The models: Claude Sonnet 5.5 (Anthropic), Gemini 3.8 Flash (Google), GPT-5.6 Sol (OpenAI) and Grok 4.7 (xAI), the general chat model each company offered in September 2026, queried through their public programming interfaces.
Because AI is trained on average. A model learns from the average of everything ever written, then gets tuned to follow instructions and please whoever is asking. Ask it for an ad and it hands back the most likely ad, and the most likely ad is the one everybody else is already running. Nobody trained it on taste, craft, or the nerve to throw out the obvious idea. That is the part a Legend brings.
Our test bears it out. Four models from four different companies kept the brief’s own fact 98% of the time. And on 21 of 50 briefs, at least two of them wrote essentially the same headline.
Then we changed the ask, and most of the sameness went away. The models do exactly what they are told, so the job is knowing what to tell them, and knowing which of the answers is any good. That part is taste.
98% of ads used the proof point we put in the brief: the 48 hours, the flat price, the 11 days. We expected models to drop the facts and reach for adjectives. We were wrong, and our pre-registered hypothesis says so below.
The models are not lazy and they are not careless. Tell them a fact and it goes in the ad. That is the good news, and it is also the trouble. A model that does exactly what the brief says will never do what the brief did not think of.
66% of headlines were made mostly of words from the brief itself. On average, 57% of a headline’s words came straight from what the owner typed.
The model’s first move is to hand your own sentence back to you, polished. That is a summary, not an idea. Only 1% of headlines asked the reader a question, and almost none tried anything the owner had not already said.
On 21 of 50 briefs, at least two different models wrote essentially the same headline (at least 70% of the same words). On 7, three of them did. On 2, all four did.
Take away the brief’s own words and the models still reach for the same language: ads written by different models for the same business shared 8.1 times as much of their added vocabulary as ads for unrelated businesses.
The briefs where the four models agreed most, picked by the code that measured it, not by hand. First try from each model, as written.
52% of ads used the long dash. 21% used an exclamation mark and 30% used an emoji. 10% of ads opened with the same two words, “Tired of”, for 16 different businesses.
We ran the same 50 briefs through the same four models again, once each, with a better ask: the ban list every Legend works under, plus one line asking for an idea the owner did not already say.
Ads with a tell fell from 71% to 2%. Headlines that were mostly the brief fell from 68% to 2%. The five most common buttons fell from 63% of all buttons to 19%, and the briefs where two models wrote the same headline fell from 21 to 0.
It was not free. Asked for an idea the owner had not stated, the models kept the brief’s own fact 84% of the time, down from 99%. And a meter can say the new headlines are different. It cannot say which one is right. Four good ideas and no way to choose is still a brief without a creative director.
| Share of ads | Plain ask | Better ask |
|---|---|---|
| Ads with at least one tell | 71% | 2% |
| Headlines mostly the brief | 68% | 2% |
| Share of the five most common buttons | 63% | 19% |
| Ads with a long dash | 50% | 0% |
| Ads with an emoji | 32% | 0% |
| Ads that kept the brief’s fact | 99% | 84% |
The same four briefs as above, first try from each model, with the better ask.
By the Ad Legends slop meter, the average ad in the study scored 7.2 out of 100, with a median of 5. Only 5% scored as slop. 73% tripped at least one tell, but most of those were a dash or an exclamation mark.
In other words, modern AI ad copy is mostly clean, on brief, and interchangeable. The meter counts tells. It cannot count the idea that is not there. That is the part a Legend judges, and the part a study in Harvard Business Review found quietly loses ground: AI ads good enough to approve that still underperform in market.
We wrote the first four down before we asked for a single ad, and the last three before the second run, each with the bar it had to clear and the result that would prove it wrong. Three held. Four did not.
| # | Hypothesis | Bar | Measured | Result |
|---|---|---|---|---|
| H1 | Slop is the default | ≥ 50% of ads trip a tell | 73% of ads tripped a tell | Held |
| H2 | The button is interchangeable | top 5 buttons ≥ 60% of all | top five buttons: 62% | Held |
| H3 | Stock phrases repeat across businesses | ≥ 3 phrases in ads for ≥ 10 businesses | 1 phrase in 10+ businesses | Wrong |
| H4 | The proof gets dropped | < 60% keep the proof | 98% kept the proof | Wrong |
| H5 | A ban list removes the tells | ads tripping a tell fall to < 25% | 2% tripped a tell | Held |
| H6 | A prompt alone will not add the idea | headlines mostly the brief stay ≥ 40% | 2% mostly the brief | Wrong |
| H7 | The button stays generic | top-5 button share stays ≥ 50% | top five: 19% | Wrong |
Each row is one of the four models, labelled A to D in an order of their own, not the order of the list above. They differ on habits: between 48% to 93% of each model’s ads tripped a tell. They agree on almost everything else.
| Model | Average slop | Tripped a tell | Kept the proof |
|---|---|---|---|
| A | 8.5 | 79% | 99% |
| B | 5.7 | 48% | 100% |
| C | 9.1 | 93% | 94% |
| D | 5.3 | 72% | 100% |
All 599 ads from both runs, with which ask wrote them, the brief’s business, the model label, the slop score, the tells it tripped and whether it kept the proof.
Ad Legends, “Why AI Ads All Look the Same,” September 2026, adlegends.ai/why-ai-ads-look-the-same.
Because AI models are trained on the average of everything ever written and tuned to follow instructions. In our study of 400 ads from four leading models, 66% of headlines were mostly the brief reworded, and on 21 of 50 briefs two different models wrote essentially the same headline.
Yes. 98% of the ads in our study used the one proof point we gave. The problem is not that AI ignores your brief. It is that it adds nothing to it.
In our study the most common button appeared on 20% of all ads, and the five most common covered 62%, across 50 unrelated businesses.
52% of the AI-written ads in our study used one. A single dash is not proof, but a dash in most sentences is one of the clearest tells.
Claude Sonnet 5.5 (Anthropic), Gemini 3.8 Flash (Google), GPT-5.6 Sol (OpenAI) and Grok 4.7 (xAI): the general chat model each company offered in September 2026. The page labels them A to D; which is which is in the full study, available on request.
None stood out. The four differed on tells (from 48% to 93% of their ads tripped one), but all of them kept the brief’s facts, reused the same buttons, and converged on the same headlines. We report them as A to D.
Change the ask. In our second run, adding the slop meter’s ban list and asking for an idea the owner had not stated cut ads with a tell from 71% to 2%, and headlines that restated the brief from 68% to 2%. Then someone with taste has to pick the right idea.
Yes. Download the CSV and cite Ad Legends, “Why AI Ads All Look the Same,” September 2026.
Check your own ad for the tells, free. Then hand it to a Legend, a licensed digital persona of a real advertising giant, who rewrites it three ways and lets a blind judge decide.