A/B Testing
A/B testing is a method for comparing two versions of something—such as a webpage, app feature, or marketing message—to see which one performs better with users. Users are typically split into groups, with each group shown a different version, and the results are measured to inform a decision. It is commonly used to produce data-driven insights rather than relying on assumptions.
A/B testing is a controlled comparison method in which an audience is divided into groups, each exposed to a different version of a variable (for example, versions A and B of a webpage or app), so that performance against a defined outcome can be measured and compared. As described in the evidence, it is applied in user-experience research and marketing contexts to compare two versions of a product or campaign element and determine which performs better. The evidence provided does not detail the underlying statistical methodology (such as significance testing, sample sizing, or randomization procedures), so this entry is scoped to the general purpose and structure of the technique as commonly presented; note that A/B testing as used for model or algorithm evaluation may involve additional considerations not covered here.
Why it matters
A/B testing matters because it replaces assumptions with evidence. As described in the evidence, it produces hard data that helps teams make informed, effective decisions rather than relying on intuition about what users prefer. For organizations building AI-enabled products or optimizing user-facing experiences, this shift toward data-driven decision-making is a foundational discipline.
In AI governance and model risk contexts, A/B testing is often invoked as a way to compare competing versions of a system in production—for example, two versions of a webpage, feature, or algorithm-driven experience. It is important to note, however, that the evidence provided here scopes A/B testing to user-experience research and marketing use, and does not address the additional considerations that arise when the technique is applied to model or algorithm evaluation. Practitioners should not assume that a marketing-style split test satisfies the validation, monitoring, or fairness assessment expectations that may apply to models under model risk management or AI governance frameworks.
Because the evidence does not detail underlying statistical methodology such as significance testing, sample sizing, or randomization, readers should treat A/B testing as a general comparison structure rather than a complete evaluation methodology. Where results feed decisions about consequential systems, the design and interpretation of the test typically warrant additional scrutiny beyond what the term itself implies.
Who it's relevant to
Inside A/B Testing
Common questions
Answers to the questions practitioners most commonly ask about A/B Testing.