Writing

How many visitors does a Shopify A/B test actually need?

The honest number, worked out from the arithmetic rather than a rule of thumb — and what it means for a store doing under 10,000 visitors a month.

11 August 2026 · 6 min read

The most common way an A/B test fails on Shopify is not a bad hypothesis or a broken script. It is a test that never had enough traffic to answer anything, stopped after a fortnight, and got read as though it had.

The number is knowable in advance, and it is worth knowing before you build the variant rather than after.

The arithmetic

Sample size depends on three things: your current conversion rate, the smallest improvement you want to be able to detect, and how confident you want to be. Nothing about the tool changes them.

For a two-group test at a 3% conversion rate, 95% confidence and 80% power:

To detectPer groupTotal
A 20% lift13,91127,822
A 10% lift53,208106,416

Halving the effect you want to catch roughly quadruples the traffic you need. That relationship is the single most useful thing on this page: it is why “let’s also see if it moves the needle a little” is an expensive sentence.

What that means for a real store

The median Shopify store sees under 10,000 visitors a month. At that traffic, and assuming every visitor enters the test:

  • A 20% lift takes about three months to resolve.
  • A 10% lift takes about ten.

That is not a limitation of any particular app. It is the arithmetic, and every tool in this category is subject to it — including Shopify’s own native testing.

If your store is under roughly 25,000 monthly visitors, A/B testing is probably not your highest-leverage activity yet. Traffic is.

We would rather say that on a public page than take a subscription from someone who will spend three months learning it.

Four ways to need less traffic

1. Test bigger changes

A new page layout moves conversion more than a button colour, and needs a quarter of the traffic to prove it. Small tests are not cheaper — they are dramatically more expensive.

2. Test fewer groups

Every group splits the same traffic. An A/B/C test needs half again as many visitors as an A/B test to reach the same confidence, and most three-arm tests exist because nobody could choose.

3. Scope to the page that matters

Testing one landing page means only its visitors enter the test — a smaller denominator, but a cleaner one. A sitewide test that changes something only 5% of visitors ever see spends 95% of its sample on people the change could not have affected.

4. Measure revenue, not clicks

Revenue per visitor is a higher-variance metric than click-through rate, so it is not a free win — but it is the number you actually act on. A test that proves a click and never checks the till has answered a question nobody asked.

What we do about it

Zinx Signal computes the required sample from your own conversion rate and shows how long the test needs at your current traffic — so “not significant yet” comes with a date instead of a shrug. On the frequentist engine, the comparison is withheld entirely until the sample is reached, because a method whose error rate depends on not peeking should not invite you to peek.

None of that makes a small store’s test faster. It just stops the result being read before it means anything.