Jun 3, 2025 / Courtlyn Saxby

Why Most Contractor A/B Tests Prove Nothing

Home / Why Most Contractor A/B Tests Prove Nothing

Most advice on split testing assumes you have the traffic to make one work. A contractor spending a few thousand a month usually does not, and running the test anyway produces a confident answer that is wrong about as often as it is right.

The arithmetic the testing advice skips

A test compares two versions and asks whether the difference between them is real or chance. Answering that needs conversions, and it needs more of them than most people expect.

Take a campaign producing forty leads a month. Split across two ads, that is twenty and twenty. Suppose one finishes on twenty-four and the other on sixteen: a sixty-forty split that looks decisive. It is not. A difference that size, on numbers that small, turns up by chance often enough that you cannot tell it from a real effect. Toss a coin forty times and a twenty-four to sixteen result is unremarkable.

Declare a winner there and you have not learned anything. You have thrown away an ad that may have been the better one.

What actually counts as enough

As a working rule, you want conversions in the high dozens per variation before a moderate difference is readable, and more than that for a small one. Clicks are not conversions: plenty of ads accumulate hundreds of clicks and a handful of booked jobs, and the booked jobs are what you are testing for.

Two things follow. Testing small differences is out of reach for most contractor accounts. And any test you do run will need to be left alone far longer than feels comfortable, often a season rather than a fortnight.

Test differences big enough to show up

If you have limited data, you can only detect large effects. So test large things.

Two headlines that say roughly the same thing in different words will differ by a few percent at most, and a few percent is invisible at your volume. A test with a real chance of producing a readable answer changes something substantive: emergency framing against price framing, a named starting price against no price, a call-only ad against one sending traffic to a page.

The rule most testing guides give: change one variable at a time, is sound, but it is often taken to mean change one small variable. Change one big one.

Click-through rate is not what you are testing

The most common way these tests go wrong has nothing to do with sample size. An ad wins on click-through rate, gets rolled out, and the booking numbers do not move.

That happens because vaguer copy nearly always earns more clicks. It excludes nobody. The ad that named a call-out fee or a service radius lost clicks from people you could not help, which is the point, and lost the test because the test was measuring the wrong thing.

Judge on booked work, or as close to it as you can measure. If your account only counts form fills and most of your leads arrive by phone, fix that before testing anything: otherwise every test is scored on a minority of your leads. That problem, and several like it, are covered in the metrics that mislead contractors.

Test the offer, not the wording

The variable with the largest effect is usually not in the copy at all. It is what the copy is offering.

Waiving a diagnostic fee, publishing a starting price, adding financing, stating same-day availability, changing which service the campaign leads with: these change who calls and how many. Verbs and adjectives do not, at least not by an amount you can measure on a contractor’s volume.

What to do instead when the volume is not there

Not testing formally does not mean guessing.

  • Let the platform rotate assets. Responsive ads test headline and description combinations continuously, across far more data than your account alone contains. Give them several genuinely different assets rather than variations on one.
  • Test the landing page, not the ad. Every campaign’s traffic pools on the page, so it accumulates data fastest. It is also where the call is won or lost.
  • Compare like seasons, not like weeks. If you must compare periods, compare this March with last March. Comparing March with January tells you about the weather.
  • Fix what is known to be broken first. An unanswered phone, an ad pointing at a general homepage, a campaign running outside your hours. These cost more than any test will recover, and none of them need a test to identify.

Frequently asked questions

How long should I run an A/B test?

Until each version has produced enough conversions to be readable, high dozens for a moderate difference, not for a set number of days. On typical contractor volume that is usually months, which is why testing small changes is rarely worth it.

Can I test if I only get twenty leads a month?

Not formally, in any reasonable timeframe. Put that effort into the landing page, the call answering, and the offer instead, and let responsive ads rotate your creative.

One ad has a much better click-through rate. Should I keep it?

Only if it also books more work. Vaguer copy reliably earns more clicks by excluding nobody, which is how a worse ad wins on the wrong measure.

Should I test the ad or the landing page?

The page. It receives traffic from every campaign, so it gathers data fastest, and it is where the decision to call is actually made.

Where to start

Before setting up a test, count the conversions your campaign produced last month and halve it. If the number is small, the honest answer is that a test will not tell you anything, and the time is better spent on the things that do not need one.

Real Time Marketing has worked with home service contractors since 2016: plumbing, HVAC, roofing, electrical and other trades, nationwide, from our office in Bradenton, Florida. Agreements are month-to-month. See our pay-per-click management or book a strategy call.

Courtlyn Saxby

Courtlyn Saxby

Courtlyn Saxby is the President of Real Time Marketing, where she leads digital marketing strategy and business growth initiatives for clients across a variety of industries. With experience supporting both...