When the vote lies
For decades, homework marks have predicted exam marks. Not perfectly, of course, but with sufficient consistency to form the basis of an entire system of indicators: parents used them to gauge whether their children were studying, teachers to adjust the difficulty of tests, and schools to allocate pupils to classes. It was an imperfect but honest indicator – a ‘proxy’, in economic terms, that correlated fairly well with what it was intended to measure. A study published in June by David Strömberg, Victor Lei and Wu Yanhui documented the moment when that proxy broke down.
Researchers tracked 26,811 Chinese students aged between 12 and 18 for thirty months. 80 per cent used generative AI models to do their homework; 20 per cent did not. After six months, the AI users’ homework marks had risen by 18 per cent and the time taken to complete their homework had fallen from 64 to 45 minutes. However, their marks in classroom exams – which were closed-book and without access to the tool – had plummeted by 20 per cent. In entrance examinations, the penalty reached 24 per cent, and only became fully apparent after two years. Those with the best homework results had the worst exam results. The correlation had not weakened: it had reversed.
An economist would recognise this structure immediately. It is Goodhart’s law: when a measure becomes a target, it ceases to be a good measure. The homework mark has become a target that can be optimised by a machine, and at that point it ceased to measure learning. AI has not corrupted the student; it has corrupted the signal. And because the entire school system was built on that signal – mid-term assessments, school reports, career guidance, access to sixth-form colleges – the corruption spreads silently, for months, before anyone realises it. The most disturbing finding of the study is the distribution of the damage. The learning losses do not affect the weakest pupils. They affect the strongest: high-achieving pupils, the youngest pupils, and those studying the humanities. Those who had the foundation to use the difficulty of the task as cognitive training are precisely those who have lost the most from its disappearance. AI has not levelled down: it has pulled the rug out from under those who were climbing the ladder.
However, a second experiment, conducted by Zara Contractor and Germán Reyes at Middlebury College, suggests that the problem does not lie with the tool itself. University students who used the chatbot to have concepts explained to them – not to have answers written for them – learnt more than their peers without access to AI, and this advantage persisted a week later. The decisive factor is not the presence of the technology, but the type of relationship the student establishes with it: if they use it as a tutor, they learn; if they use it as a ghostwriter, they unlearn. The difference, according to the Chinese data, is evident in completion times: those who spent the same amount of time on their assignments as non-users suffered almost no penalty.
Here, the issue ceases to be a pedagogical one and becomes an institutional one. Homework is the oldest form of continuous assessment, and also the least closely monitored. For decades, it worked because the cost of cheating was high: you had to physically copy the work, ask someone for help, or make up an excuse.
