CRO & Experimentation

Nearly 1,300 Tests. The Ones That Lost Taught Us More.

Nearly 1,300 CRO tests run across enterprise Shopify brands — 400+ on Rothy’s, 240+ on Four Sigmatic, 200+ on True Botanicals, 140+ on Dollar Shave Club, 130+ on Mack Weldon, 80+ each on BRUNT and Good Ranchers.

A testing program is a decision-making system, not a test counter. We start by mapping your real customer journeys across every channel, benchmark each one against the 100+ stores we monitor to find where the actual opportunity is, then design a program against business KPIs — and tell you plainly whether it’s returning more than it costs.

Programs We’ve Run

Rothy's Four Sigmatic True Botanicals Mack Weldon BRUNT Dollar Shave Club Aventon

Most testing programs are busy rather than productive

The mechanics of running an A/B test have never been easier, which is why the constraint moved. What limits most programs now is prioritization, statistical honesty, and whether anyone is measuring the program itself.

Where programs stall

A site-wide conversion rate hiding everything

Paid social, organic search, email, and returning direct traffic convert differently and fail differently, and a blended rate conceals which is actually broken.

Test velocity treated as the metric

How many tests did we run this quarter is the wrong question. The output of a program is what you now know and what you changed because of it.

Small tests, insufficient traffic

Detecting a 2% relative lift requires far more traffic than many mid-market brands have. The way out is larger changes with larger expected effects.

Nobody has priced the program

A designer, a developer, an analyst, and six weeks against a conversion move from 3% to 3.03%. At sufficient scale that can pay. Often it doesn’t, and it’s rarely checked.

Tests invalidated by your own business

Seasonality moves the numbers, and an offer change at the top of the funnel changes what the test is measuring halfway through.

Measurement that can’t support a conclusion

Aventon’s program began by standing up GTM and GA tracking for decision-grade data, because testing couldn’t function without it.

Only the wins get reported

Most decisions come from tests that didn’t win, because a loss prevents an expensive build. We report losses quarterly, alongside the wins.

  • Experience

    18+

    Years of Unlocking Growth

  • Scale

    $3.14B+

    in GMV Migrated to Shopify

  • Partnership

    1 of 5

    Founding NA Shopify Platinum Partners

  • Precision

    94.7%

    Implementation Predictability

  • Clients

    8

    Unicorns & Counting

Nearly 1,300 CRO tests run — and the program that produced 400+ of them scaled Rothy’s from $0 to a $1B+ exit.

Five things we’ve concluded about running an experimentation program

Seventeen years and nearly 1,300 tests.

1

Start with the journey, not the hypothesis

Before we write a single test, we map the real customer journeys — every path, across every channel, as they actually happen. Then we compare those journeys against baselines drawn from the stores we monitor, which is what turns a number into a judgment.

2

Prioritize against business KPIs, not against ease

We prioritize against revenue per user, lifetime value, and subscription adoption rather than implementation cost or test count.

3

Validate qualitatively before you test quantitatively

On Mack Weldon, major features were validated through usability studies before development. When testing showed customers preferred bundling on the product page, we integrated it natively — that experience went on to generate over $2M.

4

Measure downstream of the conversion event

A test that lifts conversion and hurts retention is a loss reported as a win. Thesis’s later experiments focused specifically on subscription retention rather than acquisition.

5

A testing discipline is more than a testing tool

The tool runs the split. The discipline is staging, regression testing before release, sanity testing after, and user acceptance testing — the process we stood up on Aventon.

The work

Customer journey mapping and benchmarking

Every real path through your site, by channel and intent, measured against baselines from the stores we monitor.

Program design and hypothesis backlog

A prioritized backlog scored against business KPIs, with the reasoning documented so priorities can be argued with.

Measurement and instrumentation

Event tracking, GA4 and GTM configuration, and baseline capture — verified accurate before the first test runs.

Qualitative research

Usability studies and Baymard-certified UX research to establish what’s worth testing.

Test design and statistical rigor

Effect size expectations set against your actual traffic, sample size and duration calculated in advance, stopping rules agreed before the test starts.

Experimentation infrastructure

Testing platform implementation. Optimizely is in place on Dollar Shave Club and Gaia Herbs.

Program reporting and economics

Decisions and losses reported alongside wins, plus the program’s cost against its measured incremental return.

What an experimentation engagement commits to

CommitmentAcceptance criterion
Instrumentation firstTracking verified accurate and a baseline captured before the first test ships
Powered testsSample size and duration calculated against your real traffic, before launch
Stopping rulesAgreed in advance, in writing
PrioritizationScored against business KPIs, with the scoring visible to you
Losses reportedEvery quarter, alongside the wins
Program economicsCost of the program against measured incremental return, stated plainly
Downstream effectsRetention and repeat purchase reviewed, not just the conversion event
Site safetyRegression and sanity testing around every release

If your traffic can’t support the test, running it anyway is worse than not running it.

How a program runs

Journey mapping and instrumentation

Weeks 1–3

Real journeys mapped by channel, tracking audit and repair, per-journey conversion measured against monitored-store baselines, traffic and power analysis.

Research and backlog

Weeks 2–5

Qualitative usability research, hypothesis development, prioritization against business KPIs.

Run

Ongoing

Tests designed, built, launched, and read — with regression and sanity testing around each release.

Review

Quarterly

Decisions and losses reported, program cost against measured return, backlog re-prioritized.

Timelines: A well-powered test on a high-traffic surface reads in two to four weeks. Rothy’s ran 400+ tests across an engagement that took the brand from $0 to a $1B+ acquisition, and Four Sigmatic’s 240+ tests span six years.

What we test

Product detail pages

Yotpo Reviews on Mack Weldon’s product pages lifted conversion 25%, AOV 6%, and add-to-cart among new visitors 15%. → eCommerce Funnel Engineering

Collection and listing pages

BRUNT’s listing page changes lifted conversions 20% within eleven days.

Cart and checkout

Where intent is highest and test design needs the most care. → Accessibility + Compliance Standards Design

Discovery and quiz flows

BPN’s product quiz was rebuilt end to end: completion time down 46%, completion rate up 12%. → Visual Design

Offers and pricing presentation

Bundle construction and prepay framing usually produce larger effects than interface tests.

Mobile as its own program

BRUNT’s experiments targeted mobile first; mobile conversion rose 40.7% in under two quarters. → Composable Commerce

Anatta’s Agentic Operating System

Every program runs on one system.

Senior strategists and analysts paired with an agentic framework that absorbs the commodity execution — variant builds, QA passes, reporting assembly — so senior time goes to hypothesis quality and reading results correctly.

See how we work
10x+
Faster roadmaps
35%
Boost in launch quality
78%
More roadmap flexibility

We can raise problems or run through ideas and Anatta will say what they think, explain what’s possible, and share what they’ve done before — it’s a confidence-inspiring approach.

Riley AmbroseProduct Manager, Mack Weldon

Anatta is experienced across all major platforms, with the quality and efficiency of the work standing out across the several agencies we’ve used.

Stephen HawthornwaiteCo-Founder, Rothy’s

How to audit your own testing program

Many brands reading this already run tests. This is how to find out whether the program is working.

  1. Map your journeys before anything else. Pull conversion for each path separately.
  2. Get a baseline you can compare against. A number without a benchmark is not a finding.
  3. Count decisions, not tests. If decisions are far fewer than test count, the program is generating activity rather than knowledge.
  4. Run the power calculation retroactively. Was your last result actually detectable given your traffic?
  5. Price the program. Fully loaded cost against measured incremental revenue.
  6. Audit your tracking before trusting any of the above. If two tools disagree, resolve that first.
  7. Find your losses. If you can’t produce a list of failed tests, they weren’t captured.
  8. Check your effect sizes against your ambition. If every hypothesis predicts a 1–2% lift, the program is aimed too low.
  9. Look downstream. Check retention among the users who saw a winning variant six months ago.
  10. Ask whether the program knows what marketing is doing. Offer and media changes should be on the testing calendar.

The Testing Program Audit Framework

The worksheet version, with the power calculation and program economics models included. Name and email.

Please enter a valid email address

By submitting this form you are agreeing to our Privacy Policy and to receive marketing communications from Anatta.

Frequently asked questions

Where do you find the opportunities to test?

By mapping the real customer journeys first, then comparing each against baselines from the stores we monitor — which is what turns a number into a judgment.

How is this different from funnel engineering?

This page is the testing program; funnel engineering is the buying system the program tests. → eCommerce Funnel Engineering

Do we have enough traffic to run a testing program?

Possibly not, and it’s the first thing we calculate. The options are testing larger changes, concentrating on your highest-traffic surfaces, or investing in the buying system instead.

How many tests should we run?

It’s the wrong metric. We run fewer, larger, better-prioritized experiments and report on decisions made.

What testing tool should we use?

Less important than the discipline around it. We work with Optimizely and the other major platforms.

How do you know a test result is real?

Sample size and duration calculated before launch, stopping rules agreed in writing in advance, and instrumentation verified accurate first.

What if most of our tests lose?

Most tests lose in any honest program, and that’s the program working. We report losses quarterly alongside wins for that reason.

Tell us how many tests you ran last quarter, and how many decisions came out of them.

Talk to a Strategist