Insights / Creative / 8 min read

Twenty ads per batch, or bust.

The job moved into the creative pipeline. But volume alone does not produce evidence, and a batch makes false winners more likely, not less. Here is the arithmetic and the decision rules that survive it.

Published 3 September 2026 by CurrentAds

For a decade, skill in paid social meant audience surgery: stacked interests, lookalike percentages, exclusion pyramids. That era is over, and both platforms said so themselves. Meta's and TikTok's published advertiser guidance has converged on the same advice — go broad on targeting, supply high creative diversity, refresh often — because the delivery systems now find buyers from the creative itself: who watches, who stops scrolling, who clicks. The ad is the targeting. So the advertisers who win are the ones with the most disciplined creative pipeline, and ours runs on batches of twenty.

That is the part everyone in this industry now agrees on, and it is also where most articles stop. The uncomfortable half comes next: creative volume does not by itself produce knowledge, and a batch of twenty makes it easier to fool yourself, not harder. This piece is mostly about that second half, because a testing system without a theory of evidence is just a more expensive way to guess.

Why one ad at a time tells you nothing

Start with the practice this replaces. The traditional approach is to run an ad, watch it for a fortnight, form an opinion, replace it, and repeat. Three separate problems make that approach close to worthless, and they compound rather than add.

The first is sample size, which the next section quantifies. The second is time. Two ads run in different fortnights are not being compared with each other; they are being compared with each other plus the weather, the season, a competitor's promotion, a news cycle, a price change and the day of the week. Sequential testing confounds the creative with the calendar, and there is no way to separate them afterwards. The third is the one people miss: the delivery system is not a neutral referee. It allocates impressions toward whatever performs well early, which means an ad that got a lucky first afternoon receives more delivery and then accumulates the results that confirm the choice. Run ads sequentially and you get an underpowered comparison of a moving target with a feedback loop wired into it.

The arithmetic of noise

Conversions are counts of rare events, and counts of rare events are noisy in a way that is easy to state and easy to forget. If an ad converts at a true rate of two percent, the number of conversions you observe from a given number of clicks has a standard deviation of the square root of the click count times the rate times one minus the rate. Everything below is that formula, applied at a two percent rate. Nothing here is taken from anybody's study; you can reproduce every row with a calculator.

Clicks on the adConversions expectedOrdinary luck, one deviation
500107 to 13
1,0002016 to 24
2,5005043 to 57
5,00010090 to 110
20,000400380 to 420

Arithmetic from a stated 2 percent rate. Not a benchmark, not a third-party figure.

Read the first row and the problem is obvious. Two ads that are genuinely identical, each given 500 clicks, will routinely land at 7 and 13 conversions. That is a difference of nearly two to one, produced entirely by chance, and it is exactly the gap that gets a creative promoted in a Monday meeting. The gap narrows only as the square root of volume, which is why the fifth row needs forty times the traffic of the first to look convincing.

Put a number on what a real comparison would cost. The standard approximation for the sample needed to compare two conversion rates is about sixteen times the rate times one minus the rate, divided by the square of the difference you want to detect, per variant. At two percent, detecting a lift to 2.4 percent — a twenty percent relative improvement, which is a large win in creative — comes out near nineteen thousand clicks for each ad. Multiply that by twenty ads in a batch and the required traffic is beyond almost every advertiser reading this. That is not an argument for testing less. It is an argument for being honest about what a per-ad conversion comparison can support, which is: not a verdict.

A batch of twenty does not remove the risk of crowning noise. It multiplies it, and then hands you a chart that looks decisive.

The multiple-comparison problem, stated plainly

Here is the arithmetic nobody selling creative volume puts on a slide. Suppose every ad in a batch is, in truth, exactly as good as every other, and suppose each has a one in twenty chance of looking impressive purely by luck. The probability that at least one of them looks impressive is one minus nineteen-twentieths raised to the twentieth power — about sixty-four percent. Run a batch of twenty identical ads and you will usually find a winner. Run a batch every fortnight and you will find one every fortnight, name it, build the next batch out of its descendants, and construct an entire creative strategy on top of a coin that landed heads.

This is not an argument against batches. Batches solve the sequencing problem, they expose more of the message space at once, and they suit how delivery systems actually work. It is an argument that the batch must come with decision rules that assume noise is the default explanation, because with twenty comparisons running, it usually is.

Asymmetric rules: cheap to kill, expensive to crown

The way out is to stop treating killing and crowning as the same decision. They have completely different costs. Killing an ad that was secretly fine costs you one ad out of twenty, and you have nineteen more. Crowning an ad that was secretly ordinary costs you the next quarter, because everything downstream — the iterations, the budget, the conclusions about what your market wants — inherits the error. So the rules should be loose in one direction and strict in the other.

Killing can be mechanical, and the common rule can be derived rather than borrowed. If an ad truly performs at your target cost per acquisition, conversions arrive on average once per target-CPA of spend, and the chance of seeing none at all after spending twice that is e to the minus two, about thirteen percent; after three times, about five percent. So a rule of "watch at one times target, kill at two" discards roughly one in seven ads that were actually acceptable, and a rule of "kill at three" discards about one in twenty. Choose the multiple from your own economics — how expensive your creative is to make, and how much budget you can afford to spend proving a negative — and then write it down before launch, because in the moment every ad has a defender and the argument is never about statistics.

Crowning needs three things instead of one. First, use the metrics that accumulate sample fastest for the early decisions: impressions and views are plentiful, so hook rate and click rate are measurable long before conversion rate is, and an ad that cannot earn attention will not earn conversions later. Second, require sustained spend at acceptable cost rather than a good day — a winner is an ad that has carried real budget, not one that caught a cheap conversion on Tuesday. Third, and most importantly, require replication: the winner goes into the next batch unchanged, alongside its descendants, and has to hold up a second time. Replication is the only practical defence against the sixty-four percent problem, and it costs one slot.

One prerequisite underneath all of it. Every rule above reads conversion data, so every rule above is only as good as the tracking beneath it. Deciding creative on a conversion signal that is double-counting, missing a form event or losing the click identifier is not testing, it is elaborate guessing — the failure modes are in the broken tracking epidemic.

How a batch of twenty is built

Twenty is large enough to cover real strategic variety and small enough to produce every fortnight without an in-house studio. The construction matters far more than the count: a batch is four to five distinct concepts, each expressed in four to five variations. A concept is a genuinely different argument — a customer problem story, a price and offer play, a proof cut, a founder or expert explainer, a demonstration. A variation changes the execution of that argument: the hook and the first three seconds, the format, the aspect ratio, whether a person is on screen, the call to action.

Separating those two levels is the entire point, because it turns twenty ads into two questions answered at once rather than twenty unrelated guesses. Which argument does this market respond to, and which packaging delivers it best? It also gives the statistics a fighting chance: concept-level results pool the traffic of four or five ads, so a concept comparison has several times the sample of an ad comparison and is correspondingly less likely to be noise. Judge concepts on data. Judge individual executions on data plus judgement, and hold them lightly.

Fatigue, and why the replacement is built before it is needed

Creative fatigue is not a mystery and it does not need a statistic to explain. An audience of finite size sees an ad repeatedly. The people most inclined to respond respond early, so the remaining population is progressively less inclined, and it is being shown a progressively more familiar message. Performance therefore decays as a matter of arithmetic, and it decays faster the smaller the audience and the higher the budget — a large budget aimed at a small radius exhausts its audience quickly, which is why local advertisers feel fatigue much sooner than national ones.

Refresh on evidence rather than on a fixed calendar: rising frequency alongside falling hook rate on a creative that used to work is the signature, and it is visible well before cost per acquisition moves. But build the pipeline as though the evidence is coming, because it always is. A pipeline started in response to fatigue arrives a month after it was needed, and the month in between is the one where the account panics and starts changing bids to fix a creative problem.

The fortnightly loop

Week one, the batch launches and the kill rules run daily; by the weekend the field is usually down to a handful of survivors. Week two, the survivors carry budget while the next batch is built — and this is where most teams throw away their own data. The next twenty should not be twenty new guesses. Roughly half should be descendants of what survived: the winning hook on new bodies, the winning concept in new formats, the winning spokesperson with new scripts, plus the winner itself re-entered unchanged so replication actually happens. The other half stays genuinely new, because the current winner is already on a fatigue clock and the pipeline exists to replace it before performance sags rather than after.

The usual objection is production cost, and it dissolves under modular thinking. You are not shooting twenty commercials. You are shooting one session structured for recombination: several hooks, several bodies, several closes, cut in different orders, plus static and text-on-background treatments of the same arguments. Creator content, screen recordings and plain text ads all count, and the rough cut frequently beats the polished one. That is what a creative studio is for in a paid media engagement: not twenty ideas, one session and a naming convention.

Name every asset by concept and variation, and keep the losers. Six batches is a quarter and twenty ads a batch is a hundred and twenty tested creatives, which is only an asset if the archive is still readable in three months. The record of what failed is half of what you bought: it is the map of the arguments your market does not respond to, and it is the reason the next quarter's batches start ahead of where this quarter's did. It also feeds the other half of the equation, since the best ad in the world is still spending its click on whatever the landing page does next.

None of this is glamorous. It is factory discipline applied to persuasion, with a statistician standing at the door refusing to let you crown a coin flip. That combination — volume for coverage, asymmetric rules for honesty, replication for confidence — is what separates a creative programme that compounds from one that produces a good month, a decay curve and a panic.

One companion piece, if this one landed. A creative program that compounds is a single strand of a wider argument about why no one surface carries a business on its own, and about the order the other strands are worth building in, which is set out in why one channel caps growth.

FAQ

Questions about creative testing

Why is testing one ad at a time a waste of time?

Three reasons, and they compound. The first is sample size: at ordinary conversion rates a single ad needs far more traffic than one ad's share of a normal budget to distinguish a real difference from luck. The second is time: running ads sequentially means the comparison is confounded by everything else that changed between the two windows — season, competition, promotion, news, day of week. The third is that modern delivery systems do not split traffic evenly; they reallocate toward whatever looks good early, so the ad that got a lucky first day receives more impressions and then looks like the winner it was chosen to be. One ad at a time gives you an underpowered test of a moving target with a feedback loop attached.

How many conversions does a creative test need before a winner means anything?

More than most accounts will ever give a single ad, which is the honest answer. The standard sample-size approximation for comparing two conversion rates is roughly sixteen times p times one-minus-p, divided by the square of the difference you want to detect, per variant. At a two percent conversion rate, detecting a lift to 2.4 percent works out at something like nineteen thousand clicks per ad. That is the arithmetic, and it is why classical significance testing is largely unavailable at the ad level for most advertisers. The practical consequence is not to give up on evidence. It is to stop pretending a per-ad conversion comparison is evidence, and to move the decision onto metrics that accumulate sample far faster, then confirm the survivors with sustained spend.

If I run twenty ads, am I not just multiplying my chances of a false winner?

Yes, and anyone selling batch testing without saying so is selling you a coin-flipping machine. If each ad has a one-in-twenty chance of looking impressive by luck alone, then across twenty ads the chance that at least one does is one minus 0.95 to the power of twenty, which is about sixty-four percent. A batch will therefore almost always produce something that looks like a winner even when every ad is identical in truth. The defence is not to run fewer ads. It is to make the decision rules asymmetric — cheap to kill, expensive to crown — and to require a winner to prove itself twice by surviving into the next batch.

Where does the common rule of killing at two times target cost per acquisition come from?

It can be derived rather than borrowed, which is the reason to trust it. If an ad genuinely performs at your target cost per acquisition, conversions arrive on average once per target-CPA of spend. Treating those arrivals as random, the chance of seeing none at all after spending twice the target is e to the minus two, or about thirteen percent, and after three times the target it is about five percent. So killing at two times target wrongly kills roughly one in seven ads that were actually fine, and killing at three times wrongly kills about one in twenty. Neither number is right or wrong. Pick the one that matches how expensive your creative is to produce and how much budget you have to waste on second chances.

What should actually vary between the ads in a batch?

Vary at two levels deliberately. Between concepts, vary the argument: a problem story, a price or offer play, proof, an expert explainer, a demonstration. Within a concept, vary the execution: the hook and first three seconds, the format, the aspect ratio, the caller to action, whether a person appears. Keeping those two levels separate is the whole point of a batch, because it answers two questions at once — which argument your market responds to, and which packaging of that argument delivers it best. Twenty variations of one idea is not a test, it is a superstition with a production budget.

How often does creative need refreshing?

Refresh on evidence rather than on a calendar, but do build the pipeline as though the evidence is coming, because it always does. The mechanism is straightforward: an audience of finite size sees the same ad repeatedly, the people most likely to respond respond early, and what remains is a progressively harder population being shown a progressively more familiar message. Meta's and TikTok's own published advertiser guidance has converged on the same advice — broad targeting with high creative diversity and regular refresh. The operational version is that the replacement should already be in production before the decline appears, because a pipeline started in response to fatigue arrives a month after it was needed.

Is creative volume just an excuse to bill for more production?

It can be, and the way to tell is whether the volume is modular or duplicated. Twenty genuinely separate shoots is an expensive way to learn very little. One session structured for recombination — several hooks, several bodies, several closes, cut in different orders, plus static and text treatments of the same arguments — produces the same variety at a fraction of the cost, and the archive of what lost is half the value. If a proposal prices creative per asset rather than per session, ask how many assets the recommended testing plan actually requires and price that, because that is the number you will be paying every month.

Every figure on this page is arithmetic from assumptions stated in the text. No performance result, benchmark or third-party statistic is asserted anywhere on it.

Bring the factory to your account.

Creative batches, written kill rules and a live reporting portal are part of every engagement. The free growth plan is a competitor teardown, a tracking audit and a 90 day media plan, yours to keep either way.