COGNITIKA

Part of Evidence of Educational Game Design.

How large is the effect, really?

Education research measures the size of an effect with a single number, and the number is rarely explained before it is used. That is worth settling first, because it changes how every other figure in this series should be read.

By Ilija Bojović, Founder / Lead game designer · 16 August 2026

What the number means

An effect size of about 0.2 counts as small. About 0.5 counts as moderate. 0.8 or above counts as large. These bands come from convention, not from a law of nature, but they are the convention this entire field uses.

In plain terms: a large effect (0.8) means the average learner taught the new way did better than roughly four out of five learners taught the usual way. A moderate effect (0.5) means better than roughly two out of three. A small effect (0.2) means better than just over half, which still matters once it is applied across a whole cohort, but it is not the kind of difference visible in a single classroom.

Where the numbers land

Simulation-based training sits at the top of what this field measures. Chernikova and colleagues combined 145 studies covering more than ten thousand learners and found an overall effect of 0.85, rising to 1.06 for technical performance specifically. Both are large by any convention.

Educational games in general sit much lower. Wouters and colleagues compared serious games against conventional instruction and found an effect of 0.29 on learning and 0.36 on retention. Clark and colleagues, examining digital games across school and university learners, found 0.33. All three are real and positive. None of them is large.

The distance between 0.33 and 0.85 is roughly a factor of three, and it is the question this whole series is organised around. It is not a gap between technologies. Simulations and educational games are both digital and both interactive. The gap is in the design, which is the subject of the next article.

A number belongs to the study it came from

Clark's 0.33 was measured on school and university learners. That detail is easy to drop and easy to misuse; a figure from that population does not automatically describe what happens with working adults in professional training, even though the underlying mechanism, that design quality drives the outcome, does carry across.

The cost of dropping that detail shows up clearly in one example. Dunleavy and colleagues reviewed mobile learning for health professionals: 29 trials, 3,175 learners in total, a large review. Read quickly, it is tempting to attach that scale to every number the review reports.

The review's two headline figures rest on very different ground. The knowledge effect, 0.43, comes from 11 of those trials. The skills effect, 1.12, the more impressive number, comes from only 5 trials covering 529 participants. The bigger number rests on the smaller study. Neither figure is wrong. Reporting the 1.12 beside the review's total N without saying so would be.

The rule to keep

Whether a figure can be trusted depends on three things: the population studied, the number of studies behind the specific figure rather than the review's total, and whether the figure is the review's headline or a subgroup within it. Each has an honest answer for every figure used across this series, and stating it costs a sentence.

References

Back to articles