Cybersecurity

How Microsoft uses agents and AI to cut cybersecurity costs

MAI-Cyber-1-Flash, its first artificial intelligence model specialising in cybersecurity. Alongside this, Project Perception, a new agent-based security system, has been launched.

4' min read

Translated by AI
Versione italiana

4' min read

Translated by AI
Versione italiana

Microsoft has unveiled its first artificial intelligence model specialising in cybersecurity. It is called MAI-Cyber-1-Flash.

The name sounds like that of a minor Marvel superhero. The job, however, is quite practical: reading large amounts of code, identifying where vulnerabilities are hidden, verifying them and helping to fix them.

Loading...

Alongside this model comes Project Perception, a security system comprising teams of autonomous agents. The aim is to transform cybersecurity from a series of alerts – which are often ignored – into a continuous cycle: observe, reason, act.

The difference is the same as that between a CCTV camera and an operations centre. The former sees. The latter pieces together the clues, decides which ones are significant, and sends someone to the scene.

What is MAI-Cyber-1-Flash

MAI-Cyber-1-Flash is a compact, highly code-oriented model derived from the MAI-Thinking-1 family, developed in-house by Microsoft.

It has been trained and optimised for a narrower scope: analysing complex software and identifying vulnerabilities that are difficult to spot.

Put simply, a standard language model has been through a vast library. MAI-Cyber-1-Flash has spent far more time in the section containing technical manuals, incident reports, patches, exploits and code analyses.

Microsoft claims to be able to draw on decades of experience in security, as well as **100,000 billion signals a day** from identities, devices, networks, the cloud and applications, together with operational insights generated by **1.6 million customers**. It is this wealth of data, rather than the model’s architecture alone, that constitutes the competitive advantage claimed by the company.

However, the model is not left to roam freely through corporate systems with a digital crowbar. It is integrated into MDASH, Microsoft’s multi-agent infrastructure for identifying, validating and rectifying vulnerabilities.

How it really works

The process can be described in four steps.

Loading...

The first is context analysis. The system examines the source code, software dependencies, application configuration and the relationships between the various components. A vulnerability, in fact, rarely stands alone. It is more like a crack in a bridge: it becomes dangerous depending on its location, the load it is under and any other nearby cracks.

The second step is the investigation. Various agents put forward hypotheses about possible flaws: memory management errors, incomplete authorisation checks, unfiltered inputs, weak configurations or unexpected combinations of components.

The third step is verification. The system attempts to determine whether the vulnerability can actually be exploited. This is the crucial moment, because traditional security scanners often produce long lists of potential issues. Many of these are false positives – a bit like an alarm that goes off even when the cat walks past.

The fourth step is correction. Agents can suggest a patch, check that it does not break other functions, and prepare the necessary information for the developers.

MDASH coordinates more than 100 agents and uses different models. MAI-Cyber-1-Flash is expected to handle up to 90 per cent of the tasks. For the remaining 10 per cent of the most complex issues, a larger model is deployed, currently configured as GPT-5.4.

It’s a bit like how an A&E department works. Routine cases are dealt with quickly. The consultant is called in when the X-ray starts to look like a Picasso painting.

Why does Microsoft place so much emphasis on costs

Cybersecurity is an ongoing process. Every line of code analysed, every hypothesis tested and every agent activated consumes tokens and computing power.

Always using the largest model would be like delivering every parcel by helicopter. It works, but the cost quickly becomes a concern for the finance director.

According to Microsoft, the combination of MAI-Cyber-1-Flash and GPT-5.4 within MDASH achieves around 96 per cent on CyberGym, a benchmark that measures the ability to detect vulnerabilities in large codebases. The reported result is around 12 points higher than Mythos, representing a saving of nearly 50 per cent compared with the previous MDASH configuration, which was based on GPT-5.4, GPT-5.4 mini and GPT-5.3 Codex.

A side note written in thick marker is needed here. The 96 per cent figure refers to the entire system, not MAI-Cyber-1-Flash on its own. The engine utilises the Microsoft model, GPT-5.4, the agents, the tools and the orchestration provided by MDASH.

It is therefore a comparison of complete cars, not just of engines.

Furthermore, the data is published by Microsoft. It is significant, but in order to establish a definitive record, independent assessments and comparisons carried out using the same tools, computing budgets and access to the code will be required.

The difference with Claude Mythos 5

The key development is the shift from a chatbot that suggests a fix to a system that monitors the infrastructure, simulates attacks, assesses the risk, selects the appropriate model and proposes or applies a countermeasure. The battle for AI-based cybersecurity is therefore shifting. The first phase was: who has the most intelligent model? The second will be: who is best at linking models, data, tools, permissions and procedures?

Microsoft has a clear head start in the final stretch

Copyright reserved ©
  • Luca Tremolada

    Luca TremoladaGiornalista

    Luogo: Milano via Monte Rosa 91

    Lingue parlate: Inglese, Francese

    Argomenti: Tecnologia, scienza, finanza, startup, dati

    Premi: Premio Gabriele Lanfredini sull’informazione; Premio giornalistico State Street, categoria "Innovation"; DStars 2019, categoria journalism

Loading...

Brand connect

Loading...

Newsletter

Notizie e approfondimenti sugli avvenimenti politici, economici e finanziari.

Iscriviti