Cyber security

Cybersecurity: should we slow down AI? The real question is whether we’ll be able to control it

The security of AI agents depends more on technical controls than on processing power: privileges, access controls and isolated environments are essential for preventing incidents

 Adobe Stock

6' min read

Translated by AI
Versione italiana

6' min read

Translated by AI
Versione italiana

OpenAI is reportedly open to slowing down the development of its most advanced artificial intelligence systems, according to a report by Bloomberg. The possibility was reportedly discussed by Sam Altman with staff, and the intention is to involve other laboratories as well. However, this is not a formal decision, nor is there any indication that a specific model has been postponed.

This news comes in the wake of incidents and developments that have reignited a debate going beyond the mere power of the models: to what extent can we control an agent when we entrust it with browsers, terminals, the cloud, email, repositories and other tools that enable it to operate in the digital world?

Loading...

In August, OpenAI published an analysis of an incident that occurred during a security assessment. Agents deployed in a test environment exploited vulnerabilities and insufficient isolation boundaries, gaining access to external systems not covered by the test. The problem was not an attempt to ‘escape’, but rather that the sandbox did not effectively contain the operational capabilities assigned to the models.

In another part of the assessment, the agents discovered unauthorised channels for sharing information, results and tools. This is significant because it shows that autonomy does not depend solely on the quality of the model: it also depends on the connections, permissions, shared memory and tools that the architecture makes available to it.

This sequence of events changes the nature of the discussion. The problem is no longer simply how powerful a model is, but how confident we can be that it will continue to respect the limits we have imposed on it when we allow it to use tools such as browsers, terminals, the cloud, email and repositories, which in some way open the door to dangerous interactions with the outside world. A system that produces an incorrect response can be corrected by a human being; an agent that misinterprets an objective and possesses the necessary privileges to act, on the other hand, can turn that error into a change to a system, a communication, unauthorised access or even an incident.

The difference is substantial. An agent must not, and at present cannot, be conscious, rebellious or driven by a will of its own to create a problem: it need only be capable enough to find a solution that the designers had not anticipated. If the system has access to the network, it can search for information. If it has credentials, it can use them. If it can execute code, it can turn an ambiguous instruction into a concrete action.

Recent incidents demonstrate this. In several cases, the problem was not a model that deliberately ‘disobeyed’, but an environment that left it with unforeseen capabilities and an agent skilful enough to exploit them. The lesson for cyber security is very simple: an instruction such as ‘do not do this’ is far weaker than a technical control that makes that action impossible. This is also the limitation of much of the public debate on alignment that we are currently witnessing. Alignment is often described as the ability to ensure that a model pursues goals compatible with those of humans. But in a real-world environment, it also means determining what data it can access, what systems it can reach, what tools it can use, and which actions require authorisation.

The greater the autonomy, the more difficult this problem becomes to tackle. An agent designed to achieve a goal does not necessarily reason in the same way as an employee who is familiar with the rules, context and consequences. It seeks a solution. If it encounters an obstacle, it may look for another route, unless that route is technically blocked.

And this is where the debate on the pace of development becomes interesting. Halting or slowing down the most advanced models can ease the pressure on security and allow more time for testing, audits and regulation. But slowing down the growth of capabilities does not automatically solve the problem of a less powerful agent being connected to systems that are far too sensitive.

In other words, we may end up with a model that is less intelligent but still dangerous if we grant it excessive privileges. An agent with access to corporate email, a CRM, source code or a cloud environment can cause significant damage even without being a superintelligence.

Loading...

This is why the real issue should not simply be “how far are we from superintelligence?”. The most pressing question is far more practical: how much control can we exercise over a system that makes operational decisions autonomously? Are we capable of designing safe test environments? On this point, recent internal statements at OpenAI are significant. OpenAI and some of its executives have argued that, if necessary, laboratories should coordinate to moderate the pace of development of state-of-the-art models. On 6 September, the company outlined its objective of building an automated AI researcher capable of working under human supervision and contributing to the development of new systems. “We aim to safely build an automated AI researcher, capable of working under human supervision to drive progress in deep learning and alignment, enabling iterative improvements,” OpenAI announced.

The direction is therefore clear: to use AI to accelerate AI research itself.

L’IA può davvero sterminare l’umanità?

This creates a potential source of tension. If more capable systems are used to design, train and improve even more capable systems, the pace of progress may accelerate just as it becomes more difficult to verify each step. This does not mean that some form of uncontrolled self-improvement is already taking place, but it does mean that human control must become an integral technical feature of the process, rather than merely a promise.

Meanwhile, even within the laboratories, far more concerned voices are beginning to emerge. Researchers such as Jacob Coxon, who has worked at both OpenAI and Anthropic, are warning of a race towards ever more powerful systems without adequate safeguards. His former manager, Evan Hubinger, has stated that he believes there is a probability of more than 10 per cent that an AI could contribute to the extinction of humanity within the next decade. These are personal estimates, not scientific data, and it would be a mistake to treat them as definitive predictions.

But it would be just as wrong to dismiss them as mere doomsday scenarios. The concerns of those in the know reflect a problem that is already evident: the gap between the pace at which agents’ capabilities are growing and the maturity of the tools used to monitor them.

Cybersecurity is well aware of this problem.

No administrator would entrust a software programme with a privileged account without segmentation, logging, access control and the ability to revoke access immediately. Yet we are beginning to link AI agents to tools capable of performing precisely this sort of operation.

The answer, therefore, cannot simply be to slow things down. We need truly isolated sandboxes, access based on the principle of least privilege, temporary credentials, control over outbound connections, human approval for high-impact actions, and independent logs that the agent cannot alter. Above all, testing must take place in environments that cannot turn a laboratory error into an incident involving third parties.

The political issue has already been raised.

In the United States, calls for mandatory safety regulations for advanced systems are growing, whilst OpenAI has recently argued for the need for national safety requirements. At the same time, discussions are emerging about whether laboratories should coordinate to slow down development, with issues also arising concerning competition and antitrust.It may be that, in the end, we do not need to choose between racing ahead and pausing. We may need a third option: to continue developing more capable models, whilst increasing the level of safeguards at the same pace.

The crucial question, in fact, is not whether AI will actually be capable of destroying humanity within ten years. We do not know. The question is whether we are prepared to grant ever greater autonomy to systems that we are not yet able to control with the same precision with which we know how to build them.Technology does not decide on its own how fast this race should be. It is decided by companies, investors, governments and, ultimately, the people who choose which capabilities to make available to these systems.And when it comes to AI agents, the question we must ask is: if tomorrow it were to find a path that we had not foreseen, are we certain we would still have the means to prevent it from taking that path?

* Cyber security and intelligence expert

Copyright reserved ©
Loading...

Brand connect

Loading...

Newsletter

Notizie e approfondimenti sugli avvenimenti politici, economici e finanziari.

Iscriviti