Toss a coin ten times and it might land heads seven times. That does not mean the coin is biased. It means ten tosses are not many. Ads behave the same way. An ad that produces two leads on Monday and none on Tuesday is not necessarily failing. It may simply be small numbers.
Count events, not impressions
Impressions are cheap and plentiful, so they say little. What you want to judge is the result you care about: leads, calls, bookings. The fewer of those you have, the less any comparison can tell you. If ad A has 4 leads and ad B has 6, that gap could easily reverse next week.
Rules of thumb
These are rough working guides, not laws, and the right threshold depends on your volume and how sure you need to be.
| What you are judging | A sensible minimum before deciding |
|---|---|
| Whether an ad gets attention | Enough impressions and clicks that one person’s behaviour barely changes the click rate |
| Whether a page converts | Enough visitors that a handful of extra enquiries would not flip the picture |
| Whether one ad beats another | Enough leads on each that the gap is larger than what luck could explain, and at least a full week so that weekdays and weekends both appear |
| Whether a lead source is good | Enough leads with recorded outcomes, not just counts |
Platform learning periods
Ad platforms run a learning period after a campaign starts or changes significantly, during which delivery is less stable. Meta has tied leaving that phase to a number of optimisation events within a week, historically around fifty per ad set. Check the current figure in Ads Manager, because it changes. If your budget cannot produce that many results, use fewer, larger ad sets so that each one has a chance to learn.
Give it a full cycle
Buying behaviour follows a weekly rhythm in many businesses. Judging after two days that happen to be a weekend, or a public holiday, can mislead. A full week is a minimum for most decisions.
Judge quality later
Cost per lead can be read in days. Lead quality takes longer, because you have to wait for the outcome. If you decide on cost alone, you will favour ads that produce cheap, weak leads. Build in a delay before calling a winner, and check what the leads turned into.
Small numbers flip easily
Suppose two ads each run for a week. Ad A produces five leads and ad B seven. It is tempting to call B the winner. But if both ads were in truth identical, random variation alone could easily produce a gap like that. Now suppose the numbers were fifty and seventy. The same ratio starts to mean something. The lesson is that ratios are not evidence until the counts behind them are large enough.
Patience is a skill
The urge to act on early numbers is strong. A short habit helps: write down what you think will happen and the date you will look, and then do not look at conclusions until that date. Checking spend and delivery daily is sensible. Judging performance daily is not.
A one-page decision note
Before you launch, write four lines. What we are testing. What result would make us keep it. What result would make us stop it. When we will look. That page turns a vague hope into a proper test, and it protects you from moving the goalposts after the numbers arrive.
If your volume is thin
With very few results per week, be humble about conclusions. Run fewer tests, make bigger changes between them, and focus on things you can reason about, such as whether the page works and whether the offer is clear. Our note on when to stop a campaign shows how to set the rules before you start.