The impact of AI

Over-diagnosis: the dangerous side-effect of artificial intelligence in healthcare

Algorithms promise more accurate and earlier diagnoses. But tests, imaging and new diagnostic pathways are also on the rise. The real challenge will be teaching AI not only when to look, but also when to stop

 wladimir1804 - stock.adobe.com

6' min read

Translated by AI
Versione italiana

6' min read

Translated by AI
Versione italiana

Artificial intelligence promises to help us diagnose diseases earlier and more accurately. But there is one question we are still not asking often enough: how many tests will it make us undergo to reach that diagnosis? This is no longer just a theoretical risk. A study published in Communications Medicine assessed a large language model in selecting radiological examinations in realistic clinical scenarios, comparing its decisions with the Appropriateness Criteria of the American College of Radiology. The model frequently suggested examinations that were inconsistent with the appropriateness criteria and, above all, unnecessary imaging. This is a sign that deserves serious attention. Because artificial intelligence may not merely change the way we arrive at a diagnosis. It could alter the volume of medical care we provide in order to reach it.

Every doctor knows that ordering a test is not a neutral act. A test produces an answer, but often raises a new question as well. A result just outside the normal range suggests the need for further investigation. An inconclusive image leads to a second investigative method. An incidental finding paves the way for follow-ups, further consultations and, at times, invasive procedures. This is what we call diagnostic cascade: the diagnostic cascade. Artificial intelligence has a characteristic that represents both its extraordinary strength and its potential limitation: it can simultaneously consider a volume of data, associations and diagnostic hypotheses that is vastly greater than that which the human mind can handle. But this is precisely where the problem lies. A system capable of envisaging a vast number of diseases must be equally capable of understanding which ones are not worth investigating. Otherwise, we risk creating a form of medicine that sees everything, but no longer knows what is important to look for. Statistics have always taught us this. When we increase the number of tests carried out on people with a low pre-test probability of a disease, false positives inevitably rise. Even an excellent test loses its value when applied to the wrong population. This is Bayes’ theorem, but it is also everyday medicine.

Loading...

Today, however, the problem is taking on potentially entirely new dimensions. Increasingly sophisticated imaging, biomarkers, genomics, proteomics, metabolomics, the microbiome, wearables and continuous monitoring can generate an unprecedented volume of biological signals. AI can therefore take two opposing paths. It can become the most powerful tool we have ever had for better selecting diagnostic tests, identifying the patients for whom the likelihood of benefit is genuinely high. Or it can turn into a gigantic machine for diagnostic sensitivity: always ready to suggest yet another hypothesis, another biomarker, a CT scan, an MRI, or a new test. More diagnoses, however, do not necessarily mean better health. This is the paradox we must address before AI becomes a permanent feature of clinical pathways. The problem is particularly evident with large language models, which are gradually finding their way into clinical reasoning. Their performance in benchmark tests is impressive. But real-world medicine is not a quiz in which one must identify the correct answer. It is a process made up of probabilities, successive decisions, consequences, costs and, above all, uncertainty.

A systematic review and meta-analysis published in January 2026 in npj Digital Medicine analysed studies in which doctors and large language models worked together. In certain diagnostic and management tasks, a potential benefit of human–AI collaboration emerged, but with considerable heterogeneity and still a great deal of uncertainty regarding the transferability of the results to everyday practice. The authors themselves conclude by calling for pragmatic, multicentre trials integrated into real-world workflows. A subsequent systematic review of the evidence, published in June 2026, reached an equally important conclusion: large language models are already finding their way into clinical practice, whilst the evidence base regarding their actual clinical outcomes remains incomplete. We must therefore not confuse the ability to answer a question correctly with the ability to manage a patient.

The crucial point is uncertainty. Good medicine is not merely about recognising a disease. It is also about knowing when it is not necessary to look for one. It means understanding that a value slightly outside the normal range does not necessarily indicate a disease. That an anatomical finding does not always correspond to a clinically significant disease. That sometimes it is more appropriate to observe than to investigate immediately. Above all, it means accepting that zero risk does not exist. And it is precisely this ability to tolerate and manage uncertainty that is one of the most sophisticated aspects of clinical reasoning. An algorithm incapable of concluding that ‘the available information does not justify further investigations’ can be extraordinarily powerful and, at the same time, very unwise. There is also another sign that should give us pause for thought. In 2025, *Nature Medicine* published a study analysing over 1.7 million outputs generated by nine large language models based on A&E cases. Given the same clinical information, altering certain socio-demographic characteristics of the patient changed the recommendations made by the models. In cases presented as belonging to high-income brackets, for example, the recommendation for advanced imaging, such as CT and MRI scans, increased significantly. This is a very important finding. It means that even a seemingly simple decision — whether or not to carry out a test — can be influenced by characteristics that ought not to affect its clinical appropriateness. This is why I believe it is necessary to introduce a new parameter into the evaluation of artificial intelligence applied to healthcare. We could call it diagnostic footprint, the diagnostic footprint of an algorithm. We should no longer ask merely how accurate, sensitive or specific a system is. We should also measure how much medical intervention it generates to achieve that accuracy. How many additional tests? How many CT scans? How many MRIs? How many endoscopies? How many specialist consultations? How many biopsies? How many false positives? How many incidental findings? And, at the end of the entire cascade, how many clinically significant events are actually prevented? This could become one of the new quality indicators for healthcare AI. An algorithm that marginally increases diagnostic accuracy but leads to a much greater increase in tests and procedures is not necessarily a better algorithm. It might simply be an algorithm that looks harder.

This distinction is also crucial from an economic perspective. Resources within the National Health Service are not infinite. Healthcare professionals’ time is not infinite. CT scans, MRIs, endoscopies, specialist consultations and laboratory tests have a finite capacity. If systems used by millions of citizens were to increase the likelihood of requesting further investigations, even only slightly, the consequences could be enormous. The result would be paradoxical: introducing artificial intelligence to make the healthcare system more efficient, only to discover that it has contributed to longer waiting lists. And it is not enough simply to add a computerised decision-support system to correct the problem. A cluster-randomised clinical trial published in JAMA in 2025 evaluated a clinical decision support system designed specifically to improve the appropriateness of imaging requests. The system did not significantly reduce the number of inappropriate requests made by doctors. Technology, therefore, does not resolve the issue of appropriateness on its own. It must be designed with appropriateness in mind. And this is where the doctor’s decisive role comes to the fore.

For centuries, we have regarded a doctor’s value primarily as their ability to identify what was wrong with the patient. In artificial intelligence-driven medicine, it may become just as important to determine what is not worth investigating. The ability to know when to stop is not a failure to diagnose. It is clinical expertise. And it is a skill that we must also instil in algorithms. The next generation of medical AI should not merely say: ‘This condition is a possibility’. It should be able to add: ‘The probability is so low that continuing with further investigations is likely to do more harm than good’. It should be aware not only of the sensitivity and specificity of tests, but also of the pre-test probability, predictive value, the patient’s longitudinal history, their comorbidities, and the biological, psychological and economic costs of every decision. In other words, artificial intelligence will have to learn one of the most difficult things we learn as doctors: when to stop. This is a particularly urgent issue as Europe is building a new digital health ecosystem and is progressively introducing rules for the use of artificial intelligence in high-risk sectors. The governance of healthcare AI cannot be limited to privacy, cybersecurity, the robustness of algorithms and their transparency. It must include a deeply medical principle: appropriateness. Because we are entering an era in which the problem facing medicine will no longer be a scarcity of information. It could become an excess of it. We will have algorithms capable of identifying correlations invisible to the human eye. Increasingly sensitive biomarkers. Increasingly detailed images. Almost continuous biological monitoring. We will know more and more. But precisely for this reason, we will need even more someone – whether human or machine – capable of understanding when knowing more does not necessarily mean treating better. Artificial intelligence will be able to tell us everything a patient might have. The responsibility of medicine will remain to understand what is truly worth investigating.

* Full Professor of Internal Medicine at the Catholic University of the Sacred Heart and Scientific Director of the A. Gemelli University Polyclinic IRCCS Foundation in Rome

Copyright reserved ©
Loading...

Brand connect

Loading...

Newsletter

Notizie e approfondimenti sugli avvenimenti politici, economici e finanziari.

Iscriviti