BLOG

A/B Testing Landing Pages: When It Works, and When You Do Not Have the Traffic

Published 29 July 2026

A/B testing needs volume most landing pages never see. Below a few hundred conversions a month you will read noise as signal. Two 2026 studies on AI-written copy show what a well-powered test looks like, and they reached opposite conclusions on the same question.

A/B testing sounds like the responsible way to improve a page. Show half your visitors one version, half another, keep the winner.

It works when you have the traffic. Most landing pages do not, and running a test without the volume produces a confident answer that is wrong.

The volume problem

Detecting a small improvement needs a large sample. Moving a 4% conversion rate to 5% is a 25% relative lift, and confirming it at conventional significance takes thousands of visitors per variant.

Work backwards from your own numbers. If your page sees 500 visitors a month and converts at 4%, you get 20 conversions. Split across two variants that is 10 each. A three-conversion difference between them is noise, and you will be tempted to call it a result.

Digital Applied, running 2,000 A/B tests across five funnel categories between October 2025 and March 2026, required a 95% significance threshold and a minimum of 1,000 sessions per variant, discarding tests that ran past 30 days without hitting the floor. Around 11% of their tests were excluded on those grounds. That is a well-resourced team throwing away one test in nine for insufficient data.

What to do below the threshold

Fix the structural things first. They do not need a test because the direction is known.

We crawled 800 B2B SaaS company sites in July 2026 and scored 675. 30.7% had no call to action in the page body, excluding header, nav and footer. 42.2% showed any social proof, and 11.6% placed it before the first ask. 15.4% stated a price.

Adding a CTA to a page that lacks one does not need a split test. Neither does moving your logo wall above the fold-line of the argument, or publishing a price.

Then use micro conversions. Scroll depth to pricing, clicks per CTA, form starts against completions. These arrive in the hundreds while macro conversions arrive in single digits.

Then change one big thing at a time and watch the monthly number. Less rigorous than a split test, and honest about being less rigorous.

What a properly powered test looks like

Two studies published within six months of each other asked whether AI writes better ad copy than people. They reached opposite answers, which makes the pair more instructive than either alone.

Meguellati and colleagues ran two experiments. With 400 participants, LLM-written ads tied human-written ads on personality-tailored copy: 51.1% against 48.9%, p > 0.05. With 800 participants across four psychological persuasion principles, LLM ads won: 59.1% against 40.9%, p < 0.001, strongest on authority appeals at 63.0% and consensus at 62.5%. Even after a 21.2 percentage-point penalty for being identified as AI, 29.4% still preferred the AI version knowing its origin.

Peruta and Riby at Syracuse University, working with Ipsos, found the opposite. They paired 20 human ads created before 2021 with AI counterparts generated from identical strategic briefs, and put both in front of 3,000 US respondents scored on Ipsos's sales-validated measures. Human ads over-indexed by 11 points. AI ads under-indexed by 5. Only 13% of viewers could confidently identify the AI ads.

Both are well-powered. Both are honest. They disagree because they measured different things: Meguellati measured preference between paired options, Peruta and Riby measured performance against sales-validated benchmarks, and their AI ads had to reproduce briefs written for award-winning human campaigns.

The reconciliation that fits both results: AI performs on structured persuasion appeals and struggles where the brief needs storytelling, emotion or a point of view.

Two lessons for your own testing. A well-powered study can still be answering a narrower question than its headline suggests. And when two good studies disagree, the disagreement usually hides a difference in what was measured.

What to test, in order

Start with the biggest differences. Small tests need the most traffic to resolve.

  1. The offer. A different thing being offered outranks any wording change.
  2. The headline, as a whole rewrite rather than a word swap.
  3. The form, adding or removing fields. The median page in our sample had 0 form fields, so most pages have room in one direction or the other.
  4. CTA placement and count. Unbounce found conversion falling as links multiply: 13.5% at one link, 11.9% at two to four, 10.5% at five or more.
  5. Button colour and copy. Last, and only with real volume.

Calling a test

Set the sample size before you start, not after you look. Deciding to stop when the numbers look good is how a coin flip becomes a strategy.

Run for whole weeks. Traffic behaves differently on a Tuesday than a Sunday, and a test that runs nine days over-weights one weekday.

Accept null results. A test showing no difference told you the variable does not matter as much as you thought, which is worth knowing.

Sources

  • Meguellati, E., Civelli, S., Han, L., Bernstein, A., Sadiq, S. and Demartini, G. (2025). LLM-Generated Ads: From Personalization Parity to Persuasion Superiority. arXiv:2512.03373.
  • Peruta, A. and Riby, C. (Syracuse University, S.I. Newhouse School) with Ipsos, reported 27 May 2026. n=3,000.
  • Digital Applied, Landing Page Conversion: 2,000 Pages Tested (2026). Tests run October 2025 to March 2026.
  • Unbounce, Conversion Benchmark Report. Approximately 41,000 landing pages, 464M pageviews, 57M conversions.
  • LeadzLander, The State of SaaS Landing Pages (2026). 675 pages scored, collected 2026-07-25, seed 20260725.