Over-diagnosis: the dangerous side-effect of artificial intelligence in healthcare
Algorithms promise more accurate and earlier diagnoses. But tests, imaging and new diagnostic pathways are also on the rise. The real challenge will be teaching AI not only when to look, but also when to stop
Artificial intelligence promises to help us diagnose diseases earlier and more accurately. But there is one question we are still not asking often enough: how many tests will it make us undergo to reach that diagnosis? This is no longer just a theoretical risk. A study published in Communications Medicine assessed a large language model in selecting radiological examinations in realistic clinical scenarios, comparing its decisions with the Appropriateness Criteria of the American College of Radiology. The model frequently suggested examinations that were inconsistent with the appropriateness criteria and, above all, unnecessary imaging. This is a sign that deserves serious attention. Because artificial intelligence may not merely change the way we arrive at a diagnosis. It could alter the volume of medical care we provide in order to reach it.
Every doctor knows that ordering a test is not a neutral act. A test produces an answer, but often raises a new question as well. A result just outside the normal range suggests the need for further investigation. An inconclusive image leads to a second investigative method. An incidental finding paves the way for follow-ups, further consultations and, at times, invasive procedures. This is what we call diagnostic cascade: the diagnostic cascade. Artificial intelligence has a characteristic that represents both its extraordinary strength and its potential limitation: it can simultaneously consider a volume of data, associations and diagnostic hypotheses that is vastly greater than that which the human mind can handle. But this is precisely where the problem lies. A system capable of envisaging a vast number of diseases must be equally capable of understanding which ones are not worth investigating. Otherwise, we risk creating a form of medicine that sees everything, but no longer knows what is important to look for. Statistics have always taught us this. When we increase the number of tests carried out on people with a low pre-test probability of a disease, false positives inevitably rise. Even an excellent test loses its value when applied to the wrong population. This is Bayes’ theorem, but it is also everyday medicine.
Today, however, the problem is taking on potentially entirely new dimensions. Increasingly sophisticated imaging, biomarkers, genomics, proteomics, metabolomics, the microbiome, wearables and continuous monitoring can generate an unprecedented volume of biological signals. AI can therefore take two opposing paths. It can become the most powerful tool we have ever had for better selecting diagnostic tests, identifying the patients for whom the likelihood of benefit is genuinely high. Or it can turn into a gigantic machine for diagnostic sensitivity: always ready to suggest yet another hypothesis, another biomarker, a CT scan, an MRI, or a new test. More diagnoses, however, do not necessarily mean better health. This is the paradox we must address before AI becomes a permanent feature of clinical pathways. The problem is particularly evident with large language models, which are gradually finding their way into clinical reasoning. Their performance in benchmark tests is impressive. But real-world medicine is not a quiz in which one must identify the correct answer. It is a process made up of probabilities, successive decisions, consequences, costs and, above all, uncertainty.
A systematic review and meta-analysis published in January 2026 in npj Digital Medicine analysed studies in which doctors and large language models worked together. In certain diagnostic and management tasks, a potential benefit of human–AI collaboration emerged, but with considerable heterogeneity and still a great deal of uncertainty regarding the transferability of the results to everyday practice. The authors themselves conclude by calling for pragmatic, multicentre trials integrated into real-world workflows. A subsequent systematic review of the evidence, published in June 2026, reached an equally important conclusion: large language models are already finding their way into clinical practice, whilst the evidence base regarding their actual clinical outcomes remains incomplete. We must therefore not confuse the ability to answer a question correctly with the ability to manage a patient.
The crucial point is uncertainty. Good medicine is not merely about recognising a disease. It is also about knowing when it is not necessary to look for one. It means understanding that a value slightly outside the normal range does not necessarily indicate a disease. That an anatomical finding does not always correspond to a clinically significant disease. That sometimes it is more appropriate to observe than to investigate immediately. Above all, it means accepting that zero risk does not exist. And it is precisely this ability to tolerate and manage uncertainty that is one of the most sophisticated aspects of clinical reasoning. An algorithm incapable of concluding that ‘the available information does not justify further investigations’ can be extraordinarily powerful and, at the same time, very unwise. There is also another sign that should give us pause for thought. In 2025, *Nature Medicine* published a study analysing over 1.7 million outputs generated by nine large language models based on A&E cases. Given the same clinical information, altering certain socio-demographic characteristics of the patient changed the recommendations made by the models. In cases presented as belonging to high-income brackets, for example, the recommendation for advanced imaging, such as CT and MRI scans, increased significantly. This is a very important finding. It means that even a seemingly simple decision — whether or not to carry out a test — can be influenced by characteristics that ought not to affect its clinical appropriateness. This is why I believe it is necessary to introduce a new parameter into the evaluation of artificial intelligence applied to healthcare. We could call it diagnostic footprint, the diagnostic footprint of an algorithm. We should no longer ask merely how accurate, sensitive or specific a system is. We should also measure how much medical intervention it generates to achieve that accuracy. How many additional tests? How many CT scans? How many MRIs? How many endoscopies? How many specialist consultations? How many biopsies? How many false positives? How many incidental findings? And, at the end of the entire cascade, how many clinically significant events are actually prevented? This could become one of the new quality indicators for healthcare AI. An algorithm that marginally increases diagnostic accuracy but leads to a much greater increase in tests and procedures is not necessarily a better algorithm. It might simply be an algorithm that looks harder.

