CRO & Experimentation
Nearly 1,300 Tests. The Ones That Lost Taught Us More.
Nearly 1,300 CRO tests run across enterprise Shopify brands — 400+ on Rothy’s, 240+ on Four Sigmatic, 200+ on True Botanicals, 140+ on Dollar Shave Club, 130+ on Mack Weldon, 80+ each on BRUNT and Good Ranchers.
A testing program is a decision-making system, not a test counter. We start by mapping your real customer journeys across every channel, benchmark each one against the 100+ stores we monitor to find where the actual opportunity is, then design a program against business KPIs — and tell you plainly whether it’s returning more than it costs.
Most testing programs are busy rather than productive
The mechanics of running an A/B test have never been easier, which is why the constraint moved. What limits most programs now is prioritization, statistical honesty, and whether anyone is measuring the program itself.
Where programs stall
A site-wide conversion rate hiding everything
Paid social, organic search, email, and returning direct traffic convert differently and fail differently, and a blended rate conceals which is actually broken.
Test velocity treated as the metric
How many tests did we run this quarter is the wrong question. The output of a program is what you now know and what you changed because of it.
Small tests, insufficient traffic
Detecting a 2% relative lift requires far more traffic than many mid-market brands have. The way out is larger changes with larger expected effects.
Nobody has priced the program
A designer, a developer, an analyst, and six weeks against a conversion move from 3% to 3.03%. At sufficient scale that can pay. Often it doesn’t, and it’s rarely checked.
Tests invalidated by your own business
Seasonality moves the numbers, and an offer change at the top of the funnel changes what the test is measuring halfway through.
Measurement that can’t support a conclusion
Aventon’s program began by standing up GTM and GA tracking for decision-grade data, because testing couldn’t function without it.
Only the wins get reported
Most decisions come from tests that didn’t win, because a loss prevents an expensive build. We report losses quarterly, alongside the wins.
Experience
18+
Years of Unlocking Growth
Scale
$3.14B+
in GMV Migrated to Shopify
Partnership
1 of 5
Founding NA Shopify Platinum Partners
Precision
94.7%
Implementation Predictability
Clients
8
Unicorns & Counting
Nearly 1,300 CRO tests run — and the program that produced 400+ of them scaled Rothy’s from $0 to a $1B+ exit.
Five things we’ve concluded about running an experimentation program
Seventeen years and nearly 1,300 tests.
Start with the journey, not the hypothesis
Before we write a single test, we map the real customer journeys — every path, across every channel, as they actually happen. Then we compare those journeys against baselines drawn from the stores we monitor, which is what turns a number into a judgment.
Prioritize against business KPIs, not against ease
We prioritize against revenue per user, lifetime value, and subscription adoption rather than implementation cost or test count.
Validate qualitatively before you test quantitatively
On Mack Weldon, major features were validated through usability studies before development. When testing showed customers preferred bundling on the product page, we integrated it natively — that experience went on to generate over $2M.
Measure downstream of the conversion event
A test that lifts conversion and hurts retention is a loss reported as a win. Thesis’s later experiments focused specifically on subscription retention rather than acquisition.
A testing discipline is more than a testing tool
The tool runs the split. The discipline is staging, regression testing before release, sanity testing after, and user acceptance testing — the process we stood up on Aventon.
The work
Customer journey mapping and benchmarking
Every real path through your site, by channel and intent, measured against baselines from the stores we monitor.
Program design and hypothesis backlog
A prioritized backlog scored against business KPIs, with the reasoning documented so priorities can be argued with.
Measurement and instrumentation
Event tracking, GA4 and GTM configuration, and baseline capture — verified accurate before the first test runs.
Qualitative research
Usability studies and Baymard-certified UX research to establish what’s worth testing.
Test design and statistical rigor
Effect size expectations set against your actual traffic, sample size and duration calculated in advance, stopping rules agreed before the test starts.
Experimentation infrastructure
Testing platform implementation. Optimizely is in place on Dollar Shave Club and Gaia Herbs.
Program reporting and economics
Decisions and losses reported alongside wins, plus the program’s cost against its measured incremental return.
What an experimentation engagement commits to
| Commitment | Acceptance criterion |
|---|---|
| Instrumentation first | Tracking verified accurate and a baseline captured before the first test ships |
| Powered tests | Sample size and duration calculated against your real traffic, before launch |
| Stopping rules | Agreed in advance, in writing |
| Prioritization | Scored against business KPIs, with the scoring visible to you |
| Losses reported | Every quarter, alongside the wins |
| Program economics | Cost of the program against measured incremental return, stated plainly |
| Downstream effects | Retention and repeat purchase reviewed, not just the conversion event |
| Site safety | Regression and sanity testing around every release |
If your traffic can’t support the test, running it anyway is worse than not running it.
How a program runs
Journey mapping and instrumentation
Weeks 1–3
Real journeys mapped by channel, tracking audit and repair, per-journey conversion measured against monitored-store baselines, traffic and power analysis.
Research and backlog
Weeks 2–5
Qualitative usability research, hypothesis development, prioritization against business KPIs.
Run
Ongoing
Tests designed, built, launched, and read — with regression and sanity testing around each release.
Review
Quarterly
Decisions and losses reported, program cost against measured return, backlog re-prioritized.
Timelines: A well-powered test on a high-traffic surface reads in two to four weeks. Rothy’s ran 400+ tests across an engagement that took the brand from $0 to a $1B+ acquisition, and Four Sigmatic’s 240+ tests span six years.
What we test
Product detail pages
Yotpo Reviews on Mack Weldon’s product pages lifted conversion 25%, AOV 6%, and add-to-cart among new visitors 15%. → eCommerce Funnel Engineering
Collection and listing pages
BRUNT’s listing page changes lifted conversions 20% within eleven days.
Cart and checkout
Where intent is highest and test design needs the most care. → Accessibility + Compliance Standards Design
Discovery and quiz flows
BPN’s product quiz was rebuilt end to end: completion time down 46%, completion rate up 12%. → Visual Design
Offers and pricing presentation
Bundle construction and prepay framing usually produce larger effects than interface tests.
Mobile as its own program
BRUNT’s experiments targeted mobile first; mobile conversion rose 40.7% in under two quarters. → Composable Commerce
Anatta’s Agentic Operating System
Every program runs on one system.
Senior strategists and analysts paired with an agentic framework that absorbs the commodity execution — variant builds, QA passes, reporting assembly — so senior time goes to hypothesis quality and reading results correctly.
See how we workPrograms we’ve run
Rothy’s
400+ CRO tests across an engagement that took the brand from $0 to a $1B+ exit
Explore the Rothy’s success storyFour Sigmatic
240+ CRO tests across six years, alongside a 75% improvement in reporting accuracy
Explore the Four Sigmatic success storyTrue Botanicals
200+ A/B tests, with conversion moving from 4.1% to 6.2%
Explore the True Botanicals success storyMack Weldon
130+ A/B tests across navigation, product pages, cart, and checkout, backed by qualitative usability studies
Explore the Mack Weldon success storyBRUNT
80+ CRO tests, with mobile conversion up 40.7% in under two quarters
Explore the BRUNT success storyDollar Shave Club
140+ tests, after performance engineering took throughput from one test a month to four
Explore the Dollar Shave Club success storyWe can raise problems or run through ideas and Anatta will say what they think, explain what’s possible, and share what they’ve done before — it’s a confidence-inspiring approach.
Riley AmbroseProduct Manager, Mack Weldon
Anatta is experienced across all major platforms, with the quality and efficiency of the work standing out across the several agencies we’ve used.
Stephen HawthornwaiteCo-Founder, Rothy’s
How to audit your own testing program
Many brands reading this already run tests. This is how to find out whether the program is working.
- Map your journeys before anything else. Pull conversion for each path separately.
- Get a baseline you can compare against. A number without a benchmark is not a finding.
- Count decisions, not tests. If decisions are far fewer than test count, the program is generating activity rather than knowledge.
- Run the power calculation retroactively. Was your last result actually detectable given your traffic?
- Price the program. Fully loaded cost against measured incremental revenue.
- Audit your tracking before trusting any of the above. If two tools disagree, resolve that first.
- Find your losses. If you can’t produce a list of failed tests, they weren’t captured.
- Check your effect sizes against your ambition. If every hypothesis predicts a 1–2% lift, the program is aimed too low.
- Look downstream. Check retention among the users who saw a winning variant six months ago.
- Ask whether the program knows what marketing is doing. Offer and media changes should be on the testing calendar.
The Testing Program Audit Framework
The worksheet version, with the power calculation and program economics models included. Name and email.
Frequently asked questions
Where do you find the opportunities to test?
By mapping the real customer journeys first, then comparing each against baselines from the stores we monitor — which is what turns a number into a judgment.
How is this different from funnel engineering?
This page is the testing program; funnel engineering is the buying system the program tests. → eCommerce Funnel Engineering
Do we have enough traffic to run a testing program?
Possibly not, and it’s the first thing we calculate. The options are testing larger changes, concentrating on your highest-traffic surfaces, or investing in the buying system instead.
How many tests should we run?
It’s the wrong metric. We run fewer, larger, better-prioritized experiments and report on decisions made.
What testing tool should we use?
Less important than the discipline around it. We work with Optimizely and the other major platforms.
How do you know a test result is real?
Sample size and duration calculated before launch, stopping rules agreed in writing in advance, and instrumentation verified accurate first.
What if most of our tests lose?
Most tests lose in any honest program, and that’s the program working. We report losses quarterly alongside wins for that reason.








