Sep 28

The EdTech Evidence Board has identified the right problem. Now we need faster causal evidence too

When the Department for Education announced funding for an EdTech Evidence Board, a key question was what it would mean for a product to be described as effective. The pilot report, published this month, gives a thoughtful answer. It does not treat effectiveness as a single research question. The Board asks how a provider uses research and evidence, whether a product meets a genuine educational need, whether its design is informed by evidence, whether teachers and pupils can use it successfully, whether it supports teaching and learning in practice, and whether there is evidence of positive outcomes.

A positive effect size tells schools only part of what they need to know. If a product rests on weak pedagogy, doesn’t fit into real lessons or loses pupils’ interest after a week, any benefit may never reach their classrooms. Equally, a product can be well designed, popular with teachers and enjoyable for pupils and still make little difference to learning. With curriculum time under constant pressure, schools also need to know whether using a product improves the outcomes it claims to improve. These are different questions, and they need different kinds of evidence working together.
.
Methods should hold hands
The report moves away from treating any one research design as the marker of evidence quality. It argues that qualitative and mixed-method studies can show how a product is implemented, how users experience it and the conditions in which it is likely to work. That is sensible. If teachers can’t fit a platform into lessons, or it only works in conditions few schools can provide, an effect found under ideal circumstances has limited practical value.

The report frames this partly as a tension between internal validity (did the product cause the change?) and external validity (would it work elsewhere?). We don’t think schools have to choose between these. Qualitative research, implementation studies and user evidence explain how a product is used, why teachers and pupils respond as they do, and where it might succeed. Experimental evidence answers a narrower question: did using the product lead to better outcomes than pupils would otherwise have achieved? For that question, random allocation is still one of the most powerful tools we have.

Replication links the two. One trial tells you whether a product made a difference for particular pupils in a particular setting.
Repeating randomised trials across schools, cohorts and product versions shows whether the effect holds more widely, and qualitative evidence helps explain why it varies. An EdTech evidence system needs to establish both pedagogical and causal effectiveness, because a product that improves test scores but that teachers can’t use is of little value, and so is a well-designed product that doesn’t improve learning.

Evidence that keeps pace

Timing is one of the biggest challenges for EdTech evidence. One participant in the pilot pointed out that research can be out of date by the time it is published, and the report concludes that EdTech evidence needs to become dynamic and cumulative. For AI this is especially pressing. A conventional evaluation can take years from design to publication, and in that time an AI tutor may change its underlying model, interface, feedback or content. A rigorous answer about version 2.0 is of limited use to schools already using version 5.0. The answer does not have to be less rigorous evidence. It may be faster rigorous evidence.

Randomised trials do not have to take years

Our approach is to help teachers run small randomised trials in their own classrooms, which we call micro-RCTs. Pupils take a baseline test, and the WhatWorked Teachers platform then randomises them (class level or individual), collects the anonymised scores and runs the analysis. Teachers don't choose who receives the intervention, which protects the comparison, and they don't need statistical expertise to take part.

“The shorter timeframes for data analysis mean schools don't have to wait months or years to understand whether an approach is having an impact, allowing effective interventions to be refined and rolled out more quickly.” (Sarah Williams, Primary Teaching and Learning Consultant, Gateshead Council)

A single micro-RCT is a modest piece of evidence on its own. Its value comes from being repeatable: another cohort next term, another school, or the next version of the product. The first trial gives a provisional signal, later ones test whether it holds, and over time the results show both the typical size of the effect and how much it varies.

Because these trials take place in ordinary classrooms, with teachers delivering and pupils using the product as they normally would, they test whether a product works under real conditions. That addresses the concern the report raises about effects found only in ideal circumstances
.
EdTech suppliers told the Board that cost and the length of the research process were their most frequent barriers to gathering evidence, and access to schools was raised as a further difficulty. Meanwhile, 77% of educators who responded to the survey said they would be happy to consider trialling products in their setting. Teacher-led trials are one practical way of bringing those two findings together.

An example with an AI tutoring platform
 
Our most recent study evaluated Medly, an AI tutoring platform, in GCSE science. Teachers in English secondary schools ran 39 trials across Biology, Chemistry and Physics, and 929 pupils in Years 9 and 10 took a baseline test before being randomly allocated. Half revised with Medly for four weeks, while the rest revised as they normally would, in some cases using other online platforms. The comparison was therefore with ordinary revision, not with doing nothing. The whole cycle, from set-up to report, took around twelve weeks
.
Pupils allocated to Medly scored higher on the post-test, with an overall effect size of 0.33 (95% CI 0.18 to 0.48). The effect was positive in each subject:
The result is encouraging, and some features of the design need explaining. Micro-RCTs use curriculum-aligned assessments as proxy measures of attainment, not standardised tests. Standardised assessments are usually expensive and are not sensitive enough to detect change over a few weeks of work on a single topic. The four-week intervention reflected the length of the topics being revised; micro-RCTs typically run for four to eight weeks.

Almost a third of pupils did not complete the post-test, slightly more in the control group. The trials ran in the second half of the summer term, and post-tests often clashed with end-of-term activities. Medly commissioned the evaluation and provided access to schools and usage data; WhatWorked Education carried out the evaluation independently. The full results are reported in Harrison et al. (2026), Evaluating AI Tutoring at the Speed of Innovation (https://doi.org/10.48550/arXiv.2609.14789).

The same study also produced engagement and implementation evidence. Within the Medly group, pupils who answered more questions tended to score higher: each additional question was associated with about 0.18 extra marks, after allowing for baseline attainment. Because engagement was not randomised, this is an association, and more motivated pupils may simply have done more. Alongside the randomised result, though, it suggests that how much pupils use the product matters.

Teachers who returned our process survey described different ways of using Medly. Some set it as homework and others used it in lessons, and some had attended training while others had not. They also reported login difficulties on mobile phones, and some pupils who wanted answers instead of working through the tutor’s prompts. Only six of the 39 trials returned the survey, so these findings are indicative, but they are practical things a developer can change and a subsequent trial can test.

So a single four-week study gave us a causal estimate, engagement data and a set of implementation questions to take into the next round. The most useful next step now is for other teachers, schools and cohorts to test the finding again, including on later versions of Medly.

The next step

The Board has identified the right problem in that evidence about technology can’t rest on one large study published years after a product was built, and effectiveness can’t be reduced to an effect size that ignores whether teachers and pupils can use the product. The constraints that once made randomised evidence slow and expensive are starting to ease. If teachers can run short randomised trials as part of ordinary classroom practice, causal evidence can be gathered continuously.

Alongside qualitative, implementation and pedagogical evidence, that would tell schools whether a product works, how and where it works, and whether it keeps working as the technology changes. We think that is close to the dynamic, cumulative evidence system the Board wants to build, and we would welcome conversations with schools, providers and the Board about how micro-RCTs could contribute.

References
 
Chedzey, K., Evans, S., & Lindroos Cermakova, A. (2026). Piloting the EdTech Evidence Board: Report on key findings from phase one and phase two. Department for Education / Chartered College of Teaching.

Harrison, W., Khowaja, R., Dobson, E., Uwimpuhwe, G., & Higgins, S. (2026). Evaluating AI Tutoring at the Speed of Innovation: Practitioner-Led Micro-Randomised Trials of an AI Tutoring Platform in GCSE Science. arXiv:2609.14789. https://doi.org/10.48550/arXiv.2609.14789

Created with