A/B testing on Shopify works without expensive tools if you have enough traffic. Here's how to test the right things, for the right duration, without wasting a subscription.


Guilhem Teyssier
Founder & CEO
One in seven A/B tests wins. That's the honest industry average, not the number vendors put on their landing pages. Most Shopify merchants run a test for four days, see a green number, and call it a win. It isn't. It's noise wearing a lab coat.
Here's the part nobody selling testing software wants you to hear: you probably don't need their software yet. Most stores under $1M a year in revenue can run a legitimate testing program with a spreadsheet, a free app, and enough patience to let a test finish. The tool was never the bottleneck. Traffic, discipline, and knowing what to test were.
Why most Shopify A/B tests fail before they start
Traffic kills more tests than bad ideas ever will. A store doing 800 orders a month cannot reliably A/B test a checkout button color. There simply aren't enough conversions to separate a real lift from random variance. The rule of thumb from optimization agencies is blunt: stay under 1,000 orders a month and you're in what one guide calls the "risk phase," where testing tools will happily show you a winner that isn't real.
You need volume before you need software. Shopify's own guidance suggests over 5,000 visitors to the page you're testing, and at least 100 conversions per variation, before a result means anything. Under that threshold, every "significant" result is a coin flip dressed up in a dashboard.
This is a mistake almost every new store makes. They install a $300/month testing app on day one, run five tests in a month, and none of them are trustworthy. Traffic first. Tools second.
The tools that don't cost you a subscription
You don't need Optimizely or VWO to start. Shopify's theme editor lets you duplicate a template, tweak one element, and split traffic between the two using a free app or even a manual UTM-based redirect for low-volume tests. Google Analytics 4, paired with a free Shopify testing app, covers the basics: page views, conversion events, revenue per visitor.
Paid tools earn their keep once you're running 20 or more tests a year. Below that volume, the extra statistical horsepower (Bayesian modeling, automated traffic allocation, heatmaps) is mostly wasted. Spend the money on research instead. One CRO agency puts it well:
It's 80% about the research and only 20% about testing.
Read your funnel drop-off points before you touch a single button color. Session recordings, checkout drop-off data, and support tickets tell you what to test far better than a tool's built-in "idea generator" ever will.
A free stack that actually works: Shopify's native theme editor for building the variant, GA4 for measuring revenue per visitor by segment, and a heatmap tool with a free tier for spotting where attention drops off before you even build a test. That combination costs nothing and answers 90% of what a $300/month platform answers, just slower and with more manual work on your end.
How long you actually need to run a test
Two to four weeks. Not two days. Not until Slack tells you the result is "significant" and you get excited and stop early, that's the fastest way to act on data that isn't real. Every test needs to run in full-week increments, because Tuesday shoppers and Saturday shoppers behave differently, and a three-day test skews toward whichever day happened to dominate.
Statistical significance matters, but it's not the finish line by itself. Aim for 95% confidence and a minimum sample per variation. Hit both, or the test isn't done, regardless of what the app's green checkmark says.
Monthly orders | Testing readiness | What to do instead |
Under 1,000 | Risk phase, avoid formal A/B tests | Make bold, high-conviction changes and measure trend over 30 days |
1,000 to 3,000 | Single-variable tests only | Test one page, one variable, for 3-4 full weeks |
3,000+ | Ready for multivariate tests | Run 2-3 concurrent tests on separate high-traffic pages |
What to actually test on a Shopify store
Test the pages with the most traffic and the most drop-off, not the pages that are easiest to edit. That usually means product pages and checkout, not your homepage hero banner.
High-impact tests for stores under $500k/year in revenue: product image order, shipping cost visibility, review placement, and the exact copy on your add-to-cart button. Skip favicon colors and footer link order. Those move nothing.
Store owners running research-driven programs see 20-40% win rates, well above the 1-in-7 industry baseline. The difference isn't luck. It's that they test what the data says is broken, not what the CEO thinks looks nicer.
Reading results without fooling yourself
Conversion rate alone can lie to you. A test can lift conversion rate while quietly tanking average order value, because it drops prices or hides upsells. Track revenue per visitor as your primary metric, always. A 12% lift in conversion rate paired with an 18% drop in AOV is not a win, it's a wash at best.
Segment your results by device before you trust them. A change that lifts mobile conversion by 9% can simultaneously hurt desktop by 4%. Blend those together and you'll ship a change that's actually net negative for half your traffic.
New versus returning customers matter too. A homepage change built to hook first-time visitors can annoy loyal repeat buyers who already know your site and just want to check out fast. Look at both segments separately before declaring a company-wide winner. A test that's flat overall but up 15% for new visitors and down 10% for returning ones isn't a draw, it's two different problems wearing one mask.
Common mistakes that quietly wreck a test
Peeking is the biggest one. Checking results daily and stopping the moment the graph turns green feels responsible. It's the opposite. Early "wins" regress toward nothing once the sample size catches up, and you'll have shipped a change based on three good days that don't represent your actual customer base.
Running too many tests at once on the same page is another. If you're testing button copy and shipping messaging simultaneously on your product page, you won't know which one moved the number when it moves. Isolate the variable. Always.
Ignoring seasonality ruins more tests than people admit. A test that ran through a holiday sale, a Black Friday spike, or a random viral TikTok moment isn't measuring your baseline anymore, it's measuring the event. Rerun it during a normal week before you trust the result.
Most ecommerce brands run somewhere between 24 and 60 tests a year once they're mature. New stores don't need that volume. Two or three well-run tests a quarter, on high-traffic pages, will teach you more than a dozen rushed ones on pages nobody visits.
What to prioritize if you're starting from zero
Fix your baseline conversion problems first. A/B testing improves a working store. It doesn't fix a broken one.
Get to 5,000 monthly visitors on the page you want to test before you run a formal experiment. Below that, make bold changes and track 30-day trends instead.
Start with checkout and product pages. They carry the most traffic and the highest cost of friction.
Pick revenue per visitor as your success metric, not raw conversion rate.
Run every test for a minimum of three full weeks in full-week blocks, no exceptions.
Pick one. Run it for a month. Then pick the next.
Frequently Asked Questions
How many visitors do I need before I can A/B test my Shopify store?
How long should an A/B test run on Shopify?
Do I need Shopify Plus to run A/B tests?

