Cybersecurity: should we slow down AI? The real question is whether we’ll be able to control it
The security of AI agents depends more on technical controls than on processing power: privileges, access controls and isolated environments are essential for preventing incidents
OpenAI is reportedly open to slowing down the development of its most advanced artificial intelligence systems, according to a report by Bloomberg. The possibility was reportedly discussed by Sam Altman with staff, and the intention is to involve other laboratories as well. However, this is not a formal decision, nor is there any indication that a specific model has been postponed.
This news comes in the wake of incidents and developments that have reignited a debate going beyond the mere power of the models: to what extent can we control an agent when we entrust it with browsers, terminals, the cloud, email, repositories and other tools that enable it to operate in the digital world?
In August, OpenAI published an analysis of an incident that occurred during a security assessment. Agents deployed in a test environment exploited vulnerabilities and insufficient isolation boundaries, gaining access to external systems not covered by the test. The problem was not an attempt to ‘escape’, but rather that the sandbox did not effectively contain the operational capabilities assigned to the models.
In another part of the assessment, the agents discovered unauthorised channels for sharing information, results and tools. This is significant because it shows that autonomy does not depend solely on the quality of the model: it also depends on the connections, permissions, shared memory and tools that the architecture makes available to it.
This sequence of events changes the nature of the discussion. The problem is no longer simply how powerful a model is, but how confident we can be that it will continue to respect the limits we have imposed on it when we allow it to use tools such as browsers, terminals, the cloud, email and repositories, which in some way open the door to dangerous interactions with the outside world. A system that produces an incorrect response can be corrected by a human being; an agent that misinterprets an objective and possesses the necessary privileges to act, on the other hand, can turn that error into a change to a system, a communication, unauthorised access or even an incident.

