- Views: 1
- Report Article
- Articles
- Marketing & Advertising
- Marketing Tips
What is an A/A test?
Posted: Jan 28, 2022
If we suspect that our A/B tests may be giving incorrect results (i.e. giving a winner when implementing the change does not show an improvement), it is necessary to analyze why this is happening. There are some very common causes of errors that we can rule out:
- We may have too little sample size to test minor changes, and the effect on the target metric is not detectable.
- Many times it is decided to end the test early, either because the expected sample size was reached, or because we believe that the variation being tested is losing against the control. This is not a recommended practice, since user behavior varies depending on the time of day, week and even month.
If, after ruling out these options, we are still looking for a cause that explains the discrepancy between the test and the implementation, it may be useful to think about running an A/A test: that is, testing two identical versions of the page against each other. In theory, we should see similar behaviors and similar conversions in both versions.
Advantages and disadvantages of A/A testing
At first it seems absurd to think of comparing a situation against itself, however taking the time to run an A/A test has some advantages:
- It allows us to make sure that our testing tool works correctly: in an A/A test, we should not have significant differences between our two variations of the situation (for example, two websites exactly the same), since they are identical. Therefore, if a winner is declared at the end of the test, this means that there is something to solve. It may be that the testing tool is not well configured, or it may not be effective at all. Or it may be that the test has not been configured correctly. Running an A/A test gives us the opportunity to check that everything is OK before we start testing.
- We can know the expected metrics of our control: by testing a situation against itself, if everything is configured correctly, we not only get a single metric, but a range of possible metrics for our control. Therefore, when we later run an A/B test, if our metrics fall within this range, we can consider that the result is not statistically significant.
- This is a way to decide on an appropriate sample size: when the sample size is too small, there may not be enough users in some segments, and this will impact the result of our test. By increasing the sample size, we decrease this effect, until we find the right size to set up the A/B test.
On the other hand, there are also some points against A/B testing:
- Resources and testing time are wasted: spending weeks doing an A/A test are weeks that could have been used to run an A/B test. And considering that an adequate time for a test is at least two buying cycles (ideally 4, i.e. 28 days), this can generate significant delays in an experimentation process.
- It can distract us from running valid A/B tests, which in the long run will be the ones that generate increased revenue.
- They require a much larger sample size than a conventional A/B test, since we are trying to see a difference that does not exist because we are comparing like with like.
- It is possible that we fall into a situation where a winner is declared, yet it is due to chance. In that case, although both tool and test are well configured, we falsely believe that there is a problem. Even if we run an A/A test, we are not exempt from falling into type I or II errors, as these are inherent to hypothesis testing.
How to run an A/A test:
If we never ran an A/B test, starting with an A/A test is ideal to familiarize ourselves with the tool of choice, as it has very little risk, and shows us the right way to proceed when starting to run A/B tests.
Any tool with the ability to run an A/B test allows you to run an A/A test: the sample is split equally between the two versions, but in both cases the user experience is identical.
What happens if an A/A test declares a winner?
Since we are comparing a situation with itself, randomness should dictate that there is no significant difference between them. But it could happen that one of the two iterations shows an increase in performance that we did not expect.
About the Author
Guillermo Correa. Blogger at Greydrive. insights, experimentation and Cro to optimize your e-commerce in an objective way. We work with major brands in Latam and Spain to create successful experimentation programs.