For a decade, skill in paid social meant audience surgery: stacked interests, lookalike percentages, exclusion pyramids. That era is over, and both platforms said so themselves. Meta's and TikTok's published advertiser guidance has converged on the same advice — go broad on targeting, supply high creative diversity, refresh often — because the delivery systems now find buyers from the creative itself: who watches, who stops scrolling, who clicks. The ad is the targeting. So the advertisers who win are the ones with the most disciplined creative pipeline, and ours runs on batches of twenty.
That is the part everyone in this industry now agrees on, and it is also where most articles stop. The uncomfortable half comes next: creative volume does not by itself produce knowledge, and a batch of twenty makes it easier to fool yourself, not harder. This piece is mostly about that second half, because a testing system without a theory of evidence is just a more expensive way to guess.
Why one ad at a time tells you nothing
Start with the practice this replaces. The traditional approach is to run an ad, watch it for a fortnight, form an opinion, replace it, and repeat. Three separate problems make that approach close to worthless, and they compound rather than add.
The first is sample size, which the next section quantifies. The second is time. Two ads run in different fortnights are not being compared with each other; they are being compared with each other plus the weather, the season, a competitor's promotion, a news cycle, a price change and the day of the week. Sequential testing confounds the creative with the calendar, and there is no way to separate them afterwards. The third is the one people miss: the delivery system is not a neutral referee. It allocates impressions toward whatever performs well early, which means an ad that got a lucky first afternoon receives more delivery and then accumulates the results that confirm the choice. Run ads sequentially and you get an underpowered comparison of a moving target with a feedback loop wired into it.
The arithmetic of noise
Conversions are counts of rare events, and counts of rare events are noisy in a way that is easy to state and easy to forget. If an ad converts at a true rate of two percent, the number of conversions you observe from a given number of clicks has a standard deviation of the square root of the click count times the rate times one minus the rate. Everything below is that formula, applied at a two percent rate. Nothing here is taken from anybody's study; you can reproduce every row with a calculator.
| Clicks on the ad | Conversions expected | Ordinary luck, one deviation |
|---|---|---|
| 500 | 10 | 7 to 13 |
| 1,000 | 20 | 16 to 24 |
| 2,500 | 50 | 43 to 57 |
| 5,000 | 100 | 90 to 110 |
| 20,000 | 400 | 380 to 420 |
Arithmetic from a stated 2 percent rate. Not a benchmark, not a third-party figure.
Read the first row and the problem is obvious. Two ads that are genuinely identical, each given 500 clicks, will routinely land at 7 and 13 conversions. That is a difference of nearly two to one, produced entirely by chance, and it is exactly the gap that gets a creative promoted in a Monday meeting. The gap narrows only as the square root of volume, which is why the fifth row needs forty times the traffic of the first to look convincing.
Put a number on what a real comparison would cost. The standard approximation for the sample needed to compare two conversion rates is about sixteen times the rate times one minus the rate, divided by the square of the difference you want to detect, per variant. At two percent, detecting a lift to 2.4 percent — a twenty percent relative improvement, which is a large win in creative — comes out near nineteen thousand clicks for each ad. Multiply that by twenty ads in a batch and the required traffic is beyond almost every advertiser reading this. That is not an argument for testing less. It is an argument for being honest about what a per-ad conversion comparison can support, which is: not a verdict.
A batch of twenty does not remove the risk of crowning noise. It multiplies it, and then hands you a chart that looks decisive.
The multiple-comparison problem, stated plainly
Here is the arithmetic nobody selling creative volume puts on a slide. Suppose every ad in a batch is, in truth, exactly as good as every other, and suppose each has a one in twenty chance of looking impressive purely by luck. The probability that at least one of them looks impressive is one minus nineteen-twentieths raised to the twentieth power — about sixty-four percent. Run a batch of twenty identical ads and you will usually find a winner. Run a batch every fortnight and you will find one every fortnight, name it, build the next batch out of its descendants, and construct an entire creative strategy on top of a coin that landed heads.
This is not an argument against batches. Batches solve the sequencing problem, they expose more of the message space at once, and they suit how delivery systems actually work. It is an argument that the batch must come with decision rules that assume noise is the default explanation, because with twenty comparisons running, it usually is.
Asymmetric rules: cheap to kill, expensive to crown
The way out is to stop treating killing and crowning as the same decision. They have completely different costs. Killing an ad that was secretly fine costs you one ad out of twenty, and you have nineteen more. Crowning an ad that was secretly ordinary costs you the next quarter, because everything downstream — the iterations, the budget, the conclusions about what your market wants — inherits the error. So the rules should be loose in one direction and strict in the other.
Killing can be mechanical, and the common rule can be derived rather than borrowed. If an ad truly performs at your target cost per acquisition, conversions arrive on average once per target-CPA of spend, and the chance of seeing none at all after spending twice that is e to the minus two, about thirteen percent; after three times, about five percent. So a rule of "watch at one times target, kill at two" discards roughly one in seven ads that were actually acceptable, and a rule of "kill at three" discards about one in twenty. Choose the multiple from your own economics — how expensive your creative is to make, and how much budget you can afford to spend proving a negative — and then write it down before launch, because in the moment every ad has a defender and the argument is never about statistics.
Crowning needs three things instead of one. First, use the metrics that accumulate sample fastest for the early decisions: impressions and views are plentiful, so hook rate and click rate are measurable long before conversion rate is, and an ad that cannot earn attention will not earn conversions later. Second, require sustained spend at acceptable cost rather than a good day — a winner is an ad that has carried real budget, not one that caught a cheap conversion on Tuesday. Third, and most importantly, require replication: the winner goes into the next batch unchanged, alongside its descendants, and has to hold up a second time. Replication is the only practical defence against the sixty-four percent problem, and it costs one slot.
One prerequisite underneath all of it. Every rule above reads conversion data, so every rule above is only as good as the tracking beneath it. Deciding creative on a conversion signal that is double-counting, missing a form event or losing the click identifier is not testing, it is elaborate guessing — the failure modes are in the broken tracking epidemic.
How a batch of twenty is built
Twenty is large enough to cover real strategic variety and small enough to produce every fortnight without an in-house studio. The construction matters far more than the count: a batch is four to five distinct concepts, each expressed in four to five variations. A concept is a genuinely different argument — a customer problem story, a price and offer play, a proof cut, a founder or expert explainer, a demonstration. A variation changes the execution of that argument: the hook and the first three seconds, the format, the aspect ratio, whether a person is on screen, the call to action.
Separating those two levels is the entire point, because it turns twenty ads into two questions answered at once rather than twenty unrelated guesses. Which argument does this market respond to, and which packaging delivers it best? It also gives the statistics a fighting chance: concept-level results pool the traffic of four or five ads, so a concept comparison has several times the sample of an ad comparison and is correspondingly less likely to be noise. Judge concepts on data. Judge individual executions on data plus judgement, and hold them lightly.
Fatigue, and why the replacement is built before it is needed
Creative fatigue is not a mystery and it does not need a statistic to explain. An audience of finite size sees an ad repeatedly. The people most inclined to respond respond early, so the remaining population is progressively less inclined, and it is being shown a progressively more familiar message. Performance therefore decays as a matter of arithmetic, and it decays faster the smaller the audience and the higher the budget — a large budget aimed at a small radius exhausts its audience quickly, which is why local advertisers feel fatigue much sooner than national ones.
Refresh on evidence rather than on a fixed calendar: rising frequency alongside falling hook rate on a creative that used to work is the signature, and it is visible well before cost per acquisition moves. But build the pipeline as though the evidence is coming, because it always is. A pipeline started in response to fatigue arrives a month after it was needed, and the month in between is the one where the account panics and starts changing bids to fix a creative problem.
The fortnightly loop
Week one, the batch launches and the kill rules run daily; by the weekend the field is usually down to a handful of survivors. Week two, the survivors carry budget while the next batch is built — and this is where most teams throw away their own data. The next twenty should not be twenty new guesses. Roughly half should be descendants of what survived: the winning hook on new bodies, the winning concept in new formats, the winning spokesperson with new scripts, plus the winner itself re-entered unchanged so replication actually happens. The other half stays genuinely new, because the current winner is already on a fatigue clock and the pipeline exists to replace it before performance sags rather than after.
The usual objection is production cost, and it dissolves under modular thinking. You are not shooting twenty commercials. You are shooting one session structured for recombination: several hooks, several bodies, several closes, cut in different orders, plus static and text-on-background treatments of the same arguments. Creator content, screen recordings and plain text ads all count, and the rough cut frequently beats the polished one. That is what a creative studio is for in a paid media engagement: not twenty ideas, one session and a naming convention.
Name every asset by concept and variation, and keep the losers. Six batches is a quarter and twenty ads a batch is a hundred and twenty tested creatives, which is only an asset if the archive is still readable in three months. The record of what failed is half of what you bought: it is the map of the arguments your market does not respond to, and it is the reason the next quarter's batches start ahead of where this quarter's did. It also feeds the other half of the equation, since the best ad in the world is still spending its click on whatever the landing page does next.
None of this is glamorous. It is factory discipline applied to persuasion, with a statistician standing at the door refusing to let you crown a coin flip. That combination — volume for coverage, asymmetric rules for honesty, replication for confidence — is what separates a creative programme that compounds from one that produces a good month, a decay curve and a panic.
One companion piece, if this one landed. A creative program that compounds is a single strand of a wider argument about why no one surface carries a business on its own, and about the order the other strands are worth building in, which is set out in why one channel caps growth.