Skip to content
In this article

Ideas and terms

Control group

The fair comparison that turns an impression into a number

Published 9 September 20262 min read

In one paragraph

A control group is a slice of the same work, done the old way, running alongside the new one. Comparing the two at the same time is what turns a result into a number rather than an impression. Without a control group, an improvement might be the AI system, or it might be something else entirely happening at the same time, such as a quieter period or a change of staff. A control group separates the two.

Why it matters

A before-and-after comparison on its own is easy to argue with. Between the "before" measurement and the "after" one, other things usually change too: seasonal demand shifts, a process gets tidied up, or a new person joins the team. If those changes happen to line up with the introduction of an AI system, they can look like its effect when they are not. A control group holds a comparison steady by running the old way and the new way over the same stretch of time, so anything unrelated to the AI system affects both sides equally.

How it works

  • Wherever possible, work is split between the old way and the new way at random, by team, by case, or by region, so neither group is quietly easier or harder than the other.
  • The two groups are measured on the same three things a baseline records: time, cost, and quality.
  • Where a clean split is not practical, a matched comparison or a careful before-and-after can still work, as long as other changes happening at the same time are accounted for rather than ignored.
  • A small comparison is often kept running after the main measurement finishes, so you can see whether the effect holds once the novelty has worn off and the system is handling ordinary volume.

What it looks like in practice

Here is the shape of a fair test, using case triage as the example. Half the incoming cases, chosen at random, are triaged the usual way; the other half go through the new system. Both groups are handled over the same period, by broadly similar teams, so anything else happening in that stretch, a busy spell or a change in the type of cases arriving, touches both sides equally. The comparison between the two groups, not a single number from the new system alone, is what tells the organisation whether the change is worth keeping.

How this connects to our work

A control group is one of the two ideas that make measuring what an AI system is worth a genuine measurement rather than an argument, alongside a properly captured baseline. We set the comparison up before anything is measured, so the result holds up when someone checks it.