Tests run long enough and clean enough to trust
Statistically sound A/B and multivariate testing programmes, not tests called early because a number looked good.
Where this usually breaks down
A test called early because the number looked good is worse than no test at all - it creates false confidence in a result that was likely noise. Teams under pressure to show wins often check a test daily and stop the moment the variant pulls ahead, without a sample size calculated in advance to say whether that lead is real or just early-run volatility.
The same problem shows up in reverse when a genuinely working variant gets killed too early because the first few days looked flat - weekly traffic patterns mean a test needs to run a full cycle before the result means anything.
What this service actually solves
Run properly, A/B testing means a sample size and minimum run time calculated before launch, a result read only once both are satisfied, and guardrail metrics checked alongside the primary goal so a win on one number cannot hide a loss somewhere else.
A/B testing here means running a properly powered test, with sample size calculated in advance and the result read only once significance and a full weekly traffic cycle are both reached.
How we run it
We calculate statistical power before a test launches, not after it's already running - that number sets both the sample size needed and the earliest date the result can be trusted. Tests run to that planned duration regardless of how the first few days look, and guardrail metrics are monitored throughout so a variant that wins on the primary goal but damages something else gets caught before rollout, not after.
Capabilities & deliverables
Test Design
- Hypothesis definition and variant scoping
- Minimum detectable effect setting
Sample Size & Power Calculation
- Pre-test statistical power calculation
- Minimum run time based on weekly traffic patterns
Multivariate & Sequential Testing
- Multivariate test design for interacting variables
- Correction methods for sequential testing
Statistical Validation
- Significance testing at test end
- Guardrail metric monitoring throughout
Tool Setup & QA
- VWO, Optimizely, and GA4 implementation
- Cross-device and cross-browser QA before launch
What's in scope, area by area
| Area | What we deliver |
|---|---|
| Test Plan | Hypothesis, variants, sample size, and minimum run time documented before launch |
| Implementation | Test build and cross-browser and cross-device QA in the chosen testing platform |
| Statistical Readout | Significance result plus guardrail metric check at test end |
| Knowledge Base | Documented result added to a searchable test history, win or lose |
How an engagement runs
Hypothesis & Test Design
Each test starts from a specific, evidence-backed hypothesis rather than a general idea worth trying.
Sample Size Calculation
Statistical power sets the sample size and minimum run time before anything is built.
Build & QA
Variants are QA-checked across devices and browsers before the test goes live.
Launch & Monitor
The test runs to its planned duration - we track it but do not call it early because an early number looks good.
Statistical Readout
Significance and guardrail metrics are both checked before declaring a winner.
Documentation & Knowledge Base
The result, including inconclusive or negative ones, is documented so the next test builds on it.
How this compares
| Properly Powered Test | Test Called Early |
|---|---|
| Sample size is set before the test starts | The test is stopped once a number looks good |
| The result holds up if the test is run again | A large share of apparent wins turn out to be noise |
| Guardrail metrics are checked, not just the primary goal | A win on one metric can hide a loss somewhere else |
Checking results daily and stopping the moment a variant pulls ahead is one of the most common ways a testing programme quietly generates false wins.
What this changes for the business
- Test results hold up when re-run, instead of being one-off noise mistaken for a win
- Guardrail metrics catch a variant that wins on the primary goal but damages something else
- A searchable history of past tests prevents re-running an already-answered question
Who needs this
Teams running tests without a sample size calculation
If a test can be stopped whenever a number looks convincing, there was never a real stopping rule in the first place.
Sites with a testing tool installed but no real programme
Having VWO or Optimizely installed is not the same as running a disciplined testing programme around it.
Related work
We're still building out published proof for this specific service — ask us directly and we'll walk through relevant examples.
Common questions
An A/B test compares two full versions of a page against each other. A multivariate test changes several elements at once and measures how they interact, which needs substantially more traffic to reach significance since it's effectively testing multiple combinations simultaneously.
Long enough to hit the pre-calculated sample size and to cover at least one full weekly traffic cycle, since behaviour on a Tuesday and a Saturday can differ enough to distort an early read.
Checking for technical problems - broken tracking, a variant rendering incorrectly - is fine and worth doing early. Checking the conversion numbers and using them to decide whether to stop the test early is the part that undermines the statistics.
It depends on your existing stack, traffic volume, and what you need beyond simple A/B splits - GA4's native experimentation is more limited than a dedicated platform like VWO or Optimizely, which matters more for multivariate or sequential testing.
No - most individual tests are inconclusive or come back negative, and that's a normal, useful outcome, not a failure. A testing programme's value comes from the accumulated, trustworthy knowledge across many tests, not from every single test winning.
Already running tests but not sure you can trust the results?
We'll review your current testing setup for sample size and stopping-rule problems before recommending anything.
Talk to us →