When the Department for Education announced funding for an EdTech Evidence Board, a key question was what it would mean for a product to be described as effective. The pilot report, published this month, gives a thoughtful answer. It does not treat effectiveness as a single research question. The Board asks how a provider uses research and evidence, whether a product meets a genuine educational need, whether its design is informed by evidence, whether teachers and pupils can use it successfully, whether it supports teaching and learning in practice, and whether there is evidence of positive outcomes.
Replication links the two. One trial tells you whether a product made a difference for particular pupils in a particular setting.
Repeating randomised trials across schools, cohorts and product versions shows whether the effect holds more widely, and qualitative evidence helps explain why it varies. An EdTech evidence system needs to establish both pedagogical and causal effectiveness, because a product that improves test scores but that teachers can’t use is of little value, and so is a well-designed product that doesn’t improve learning.
Evidence that keeps pace
Timing is one of the biggest challenges for EdTech evidence. One participant in the pilot pointed out that research can be out of date by the time it is published, and the report concludes that EdTech evidence needs to become dynamic and cumulative. For AI this is especially pressing. A conventional evaluation can take years from design to publication, and in that time an AI tutor may change its underlying model, interface, feedback or content. A rigorous answer about version 2.0 is of limited use to schools already using version 5.0. The answer does not have to be less rigorous evidence. It may be faster rigorous evidence.
Randomised trials do not have to take years
Our approach is to help teachers run small randomised trials in their own classrooms, which we call micro-RCTs. Pupils take a baseline test, and the WhatWorked Teachers platform then randomises them (class level or individual), collects the anonymised scores and runs the analysis. Teachers don't choose who receives the intervention, which protects the comparison, and they don't need statistical expertise to take part.
A positive effect size tells schools only part of what they need to know. If a product rests on weak pedagogy, doesn’t fit into real lessons or loses pupils’ interest after a week, any benefit may never reach their classrooms. Equally, a product can be well designed, popular with teachers and enjoyable for pupils and still make little difference to learning. With curriculum time under constant pressure, schools also need to know whether using a product improves the outcomes it claims to improve. These are different questions, and they need different kinds of evidence working together.
.
Methods should hold hands
The report moves away from treating any one research design as the marker of evidence quality. It argues that qualitative and mixed-method studies can show how a product is implemented, how users experience it and the conditions in which it is likely to work. That is sensible. If teachers can’t fit a platform into lessons, or it only works in conditions few schools can provide, an effect found under ideal circumstances has limited practical value.
Methods should hold hands
The report moves away from treating any one research design as the marker of evidence quality. It argues that qualitative and mixed-method studies can show how a product is implemented, how users experience it and the conditions in which it is likely to work. That is sensible. If teachers can’t fit a platform into lessons, or it only works in conditions few schools can provide, an effect found under ideal circumstances has limited practical value.
The report frames this partly as a tension between internal validity (did the product cause the change?) and external validity (would it work elsewhere?). We don’t think schools have to choose between these. Qualitative research, implementation studies and user evidence explain how a product is used, why teachers and pupils respond as they do, and where it might succeed. Experimental evidence answers a narrower question: did using the product lead to better outcomes than pupils would otherwise have achieved? For that question, random allocation is still one of the most powerful tools we have.
Replication links the two. One trial tells you whether a product made a difference for particular pupils in a particular setting.
Repeating randomised trials across schools, cohorts and product versions shows whether the effect holds more widely, and qualitative evidence helps explain why it varies. An EdTech evidence system needs to establish both pedagogical and causal effectiveness, because a product that improves test scores but that teachers can’t use is of little value, and so is a well-designed product that doesn’t improve learning.
Evidence that keeps pace
Timing is one of the biggest challenges for EdTech evidence. One participant in the pilot pointed out that research can be out of date by the time it is published, and the report concludes that EdTech evidence needs to become dynamic and cumulative. For AI this is especially pressing. A conventional evaluation can take years from design to publication, and in that time an AI tutor may change its underlying model, interface, feedback or content. A rigorous answer about version 2.0 is of limited use to schools already using version 5.0. The answer does not have to be less rigorous evidence. It may be faster rigorous evidence.
Randomised trials do not have to take years
Our approach is to help teachers run small randomised trials in their own classrooms, which we call micro-RCTs. Pupils take a baseline test, and the WhatWorked Teachers platform then randomises them (class level or individual), collects the anonymised scores and runs the analysis. Teachers don't choose who receives the intervention, which protects the comparison, and they don't need statistical expertise to take part.
“The shorter timeframes for data analysis mean schools don't have to wait months or years to understand whether an approach is having an impact, allowing effective interventions to be refined and rolled out more quickly.” (Sarah Williams, Primary Teaching and Learning Consultant, Gateshead Council)
A single micro-RCT is a modest piece of evidence on its own. Its value comes from being repeatable: another cohort next term, another school, or the next version of the product. The first trial gives a provisional signal, later ones test whether it holds, and over time the results show both the typical size of the effect and how much it varies.
Because these trials take place in ordinary classrooms, with teachers delivering and pupils using the product as they normally would, they test whether a product works under real conditions. That addresses the concern the report raises about effects found only in ideal circumstances
.
Because these trials take place in ordinary classrooms, with teachers delivering and pupils using the product as they normally would, they test whether a product works under real conditions. That addresses the concern the report raises about effects found only in ideal circumstances
.
EdTech suppliers told the Board that cost and the length of the research process were their most frequent barriers to gathering evidence, and access to schools was raised as a further difficulty. Meanwhile, 77% of educators who responded to the survey said they would be happy to consider trialling products in their setting. Teacher-led trials are one practical way of bringing those two findings together.
An example with an AI tutoring platform
Our most recent study evaluated Medly, an AI tutoring platform, in GCSE science. Teachers in English secondary schools ran 39 trials across Biology, Chemistry and Physics, and 929 pupils in Years 9 and 10 took a baseline test before being randomly allocated. Half revised with Medly for four weeks, while the rest revised as they normally would, in some cases using other online platforms. The comparison was therefore with ordinary revision, not with doing nothing. The whole cycle, from set-up to report, took around twelve weeks
.
An example with an AI tutoring platform
Our most recent study evaluated Medly, an AI tutoring platform, in GCSE science. Teachers in English secondary schools ran 39 trials across Biology, Chemistry and Physics, and 929 pupils in Years 9 and 10 took a baseline test before being randomly allocated. Half revised with Medly for four weeks, while the rest revised as they normally would, in some cases using other online platforms. The comparison was therefore with ordinary revision, not with doing nothing. The whole cycle, from set-up to report, took around twelve weeks
.
Pupils allocated to Medly scored higher on the post-test, with an overall effect size of 0.33 (95% CI 0.18 to 0.48). The effect was positive in each subject: