A/B Testing

Tests run long enough and clean enough to trust

Statistically sound A/B and multivariate testing programmes, not tests called early because a number looked good.

The challenge

Where this usually breaks down

A test called early because the number looked good is worse than no test at all - it creates false confidence in a result that was likely noise. Teams under pressure to show wins often check a test daily and stop the moment the variant pulls ahead, without a sample size calculated in advance to say whether that lead is real or just early-run volatility.

The same problem shows up in reverse when a genuinely working variant gets killed too early because the first few days looked flat - weekly traffic patterns mean a test needs to run a full cycle before the result means anything.

What we fix

What this service actually solves

Run properly, A/B testing means a sample size and minimum run time calculated before launch, a result read only once both are satisfied, and guardrail metrics checked alongside the primary goal so a win on one number cannot hide a loss somewhere else.

In simple terms

A/B testing here means running a properly powered test, with sample size calculated in advance and the result read only once significance and a full weekly traffic cycle are both reached.

Our approach

How we run it

We calculate statistical power before a test launches, not after it's already running - that number sets both the sample size needed and the earliest date the result can be trusted. Tests run to that planned duration regardless of how the first few days look, and guardrail metrics are monitored throughout so a variant that wins on the primary goal but damages something else gets caught before rollout, not after.

What's included

Capabilities & deliverables

01

Test Design

  • Hypothesis definition and variant scoping
  • Minimum detectable effect setting
02

Sample Size & Power Calculation

  • Pre-test statistical power calculation
  • Minimum run time based on weekly traffic patterns
03

Multivariate & Sequential Testing

  • Multivariate test design for interacting variables
  • Correction methods for sequential testing
04

Statistical Validation

  • Significance testing at test end
  • Guardrail metric monitoring throughout
05

Tool Setup & QA

  • VWO, Optimizely, and GA4 implementation
  • Cross-device and cross-browser QA before launch
Scope

What's in scope, area by area

AreaWhat we deliver
Test PlanHypothesis, variants, sample size, and minimum run time documented before launch
ImplementationTest build and cross-browser and cross-device QA in the chosen testing platform
Statistical ReadoutSignificance result plus guardrail metric check at test end
Knowledge BaseDocumented result added to a searchable test history, win or lose
Process

How an engagement runs

Hypothesis & Test Design

Each test starts from a specific, evidence-backed hypothesis rather than a general idea worth trying.

Sample Size Calculation

Statistical power sets the sample size and minimum run time before anything is built.

Build & QA

Variants are QA-checked across devices and browsers before the test goes live.

Launch & Monitor

The test runs to its planned duration - we track it but do not call it early because an early number looks good.

Statistical Readout

Significance and guardrail metrics are both checked before declaring a winner.

Documentation & Knowledge Base

The result, including inconclusive or negative ones, is documented so the next test builds on it.

In context

How this compares

Properly Powered TestTest Called Early
Sample size is set before the test startsThe test is stopped once a number looks good
The result holds up if the test is run againA large share of apparent wins turn out to be noise
Guardrail metrics are checked, not just the primary goalA win on one metric can hide a loss somewhere else

Checking results daily and stopping the moment a variant pulls ahead is one of the most common ways a testing programme quietly generates false wins.

Tools & technologies
VWOOptimizelyGA4Statistical Significance Calculators
Outcomes

What this changes for the business

  • Test results hold up when re-run, instead of being one-off noise mistaken for a win
  • Guardrail metrics catch a variant that wins on the primary goal but damages something else
  • A searchable history of past tests prevents re-running an already-answered question
Who this is for

Who needs this

Teams running tests without a sample size calculation

If a test can be stopped whenever a number looks convincing, there was never a real stopping rule in the first place.

Sites with a testing tool installed but no real programme

Having VWO or Optimizely installed is not the same as running a disciplined testing programme around it.

Proof

Related work

We're still building out published proof for this specific service — ask us directly and we'll walk through relevant examples.

FAQs

Common questions

An A/B test compares two full versions of a page against each other. A multivariate test changes several elements at once and measures how they interact, which needs substantially more traffic to reach significance since it's effectively testing multiple combinations simultaneously.

Long enough to hit the pre-calculated sample size and to cover at least one full weekly traffic cycle, since behaviour on a Tuesday and a Saturday can differ enough to distort an early read.

Checking for technical problems - broken tracking, a variant rendering incorrectly - is fine and worth doing early. Checking the conversion numbers and using them to decide whether to stop the test early is the part that undermines the statistics.

It depends on your existing stack, traffic volume, and what you need beyond simple A/B splits - GA4's native experimentation is more limited than a dedicated platform like VWO or Optimizely, which matters more for multivariate or sequential testing.

No - most individual tests are inconclusive or come back negative, and that's a normal, useful outcome, not a failure. A testing programme's value comes from the accumulated, trustworthy knowledge across many tests, not from every single test winning.

Already running tests but not sure you can trust the results?

We'll review your current testing setup for sample size and stopping-rule problems before recommending anything.

Talk to us →

Page Optimization Data

Primary Topic
A/B Testing
Primary Intent
commercial - service research
Suggested URL
/services/conversion-optimization/ab-testing
Breadcrumb
Home / Services / Conversion Optimization / A/B Testing
Key Entities
A/B TestingMultivariate TestingStatistical SignificanceSample SizeVWOOptimizely