Opinions

When the vote lies

3' min read

Translated by AI
Versione italiana

3' min read

Translated by AI
Versione italiana

For decades, homework marks have predicted exam marks. Not perfectly, of course, but with sufficient consistency to form the basis of an entire system of indicators: parents used them to gauge whether their children were studying, teachers to adjust the difficulty of tests, and schools to allocate pupils to classes. It was an imperfect but honest indicator – a ‘proxy’, in economic terms, that correlated fairly well with what it was intended to measure. A study published in June by David Strömberg, Victor Lei and Wu Yanhui documented the moment when that proxy broke down.

Researchers tracked 26,811 Chinese students aged between 12 and 18 for thirty months. 80 per cent used generative AI models to do their homework; 20 per cent did not. After six months, the AI users’ homework marks had risen by 18 per cent and the time taken to complete their homework had fallen from 64 to 45 minutes. However, their marks in classroom exams – which were closed-book and without access to the tool – had plummeted by 20 per cent. In entrance examinations, the penalty reached 24 per cent, and only became fully apparent after two years. Those with the best homework results had the worst exam results. The correlation had not weakened: it had reversed.

Loading...

An economist would recognise this structure immediately. It is Goodhart’s law: when a measure becomes a target, it ceases to be a good measure. The homework mark has become a target that can be optimised by a machine, and at that point it ceased to measure learning. AI has not corrupted the student; it has corrupted the signal. And because the entire school system was built on that signal – mid-term assessments, school reports, career guidance, access to sixth-form colleges – the corruption spreads silently, for months, before anyone realises it. The most disturbing finding of the study is the distribution of the damage. The learning losses do not affect the weakest pupils. They affect the strongest: high-achieving pupils, the youngest pupils, and those studying the humanities. Those who had the foundation to use the difficulty of the task as cognitive training are precisely those who have lost the most from its disappearance. AI has not levelled down: it has pulled the rug out from under those who were climbing the ladder.

However, a second experiment, conducted by Zara Contractor and Germán Reyes at Middlebury College, suggests that the problem does not lie with the tool itself. University students who used the chatbot to have concepts explained to them – not to have answers written for them – learnt more than their peers without access to AI, and this advantage persisted a week later. The decisive factor is not the presence of the technology, but the type of relationship the student establishes with it: if they use it as a tutor, they learn; if they use it as a ghostwriter, they unlearn. The difference, according to the Chinese data, is evident in completion times: those who spent the same amount of time on their assignments as non-users suffered almost no penalty.

Here, the issue ceases to be a pedagogical one and becomes an institutional one. Homework is the oldest form of continuous assessment, and also the least closely monitored. For decades, it worked because the cost of cheating was high: you had to physically copy the work, ask someone for help, or make up an excuse.

AI has reduced that cost to zero, and in doing so has hollowed out the system from within. What remains is a metric that everyone continues to monitor but which no longer tells us anything – like a thermometer that always reads 36.5 whilst the fever rises.

The lesson extends far beyond the classroom. In every sector where generative AI is used to produce assessable outputs – business reports, market analyses, legal opinions, code – the same risk exists: that the apparent quality of the product improves whilst the expertise of those producing it deteriorates. The lag between these two developments is what causes the damage, because it masks the loss at a time when it could still be recovered. A company that assesses its analysts on the basis of the reports they produce may discover, in a few years’ time, that it has been paying ever-higher salaries for ever-diminishing expertise.

The question is not whether to ban AI in schools or offices. The question is which metrics still hold water, and which have already become mere window dressing. Any institution that assesses people on the basis of outputs that can be produced by a machine has a problem with its measurement methods. If it does not address this now, it will find out the hard way when the time comes.

Copyright reserved ©
Loading...

Brand connect

Loading...

Newsletter

Notizie e approfondimenti sugli avvenimenti politici, economici e finanziari.

Iscriviti