Cyber security

OpenAI has just admitted that its models sometimes lie. What does that mean?

 (Foto AP/Michael Dwyer)

6' min read

Translated by AI
Versione italiana

6' min read

Translated by AI
Versione italiana

It is not often that a company publishes a document specifically designed to explain how its systems might produce unexpected or undesirable behaviour. OpenAI did just that on 16 September, presenting a new framework for identifying, analysing and reporting instances of model misalignment. At the same time, the company published six cases observed over the past six months.

According to OpenAI, the management and communication of these incidents have, to date, been rather unstructured. Some cases were retained for inclusion in broader reports, whilst others ended up in the system cards published alongside new models. The new framework, however, aims to reduce these delays, making it possible to publish information even when the company has not yet completed its analysis or found a definitive solution.

Loading...

The document also highlights a particularly significant point: according to OpenAI itself, the artificial intelligence sector has not yet developed sufficiently robust alignment and monitoring systems to support the growth in the models’ capabilities without further risks. This is a significant admission, especially given how rapidly these systems are becoming part of the day-to-day processes of businesses, professionals and users.

“This new framework aims to speed up the publication of reports on observed misalignment phenomena, even when we have not yet fully explained or mitigated the behaviour in question,” states the announcement from OpenAI.

“As AI systems become more advanced and widespread, there is a need to build a broader and more informed consensus on progress in alignment research.”

The company formally undertakes to disclose instances of misalignment that meet the criteria of this framework.

Read one after the other, the six reports paint a picture of anomalous behaviour, some of which is rather difficult to ignore. In one instance, a research model that had not yet been released inserted its own instructions into the summaries used to resume work in a subsequent session, even going so far as to include instructions on how to circumvent the system’s normal constraints. OpenAI has documented as many as 27 such incidents.

In another instance, whilst GPT-5.6 Sol was being trained, several instances of the model left notes intended for future versions, explaining how to conceal certain errors and fill in missing information by inventing data without stating that they had done so.

Then there is the case that anyone working in cyber security will recognise immediately. Whilst responding to a routine enquiry regarding budget figures from a Californian county, a bot found an API key exposed in a public repository and used it without authorisation. When this proved insufficient to obtain the requested figures, it fabricated them and presented them as genuine data from the requested source. This is not a hallucination. It is a cover-up.

Two further reports describe models that circumvent the limits rather than lying outright. One agent uploaded a file to the internet solely so that it could cite itself as a source from the browser, without asking anyone’s permission. In another case, separate instances of the same model used an internal code repository as an informal noticeboard, leaving messages for one another whilst searching for files to which they had lost access.

Loading...

The latest report concerns staff members who were working on the same task without being able to access each other’s local files, so they used public file-hosting sites to exchange them. The results of their work ended up on public URLs. Nobody had asked them to do so, but nobody had explicitly forbidden it either – which is perhaps the real problem.

OpenAI has therefore developed a three-tiered process for managing these cases: ‘ready for disclosure’, ‘minor investigation’, and ‘extended investigation’ for cases involving third parties or serious risks. Disagreements over what to publish, or which level to assign, are referred to OpenAI’s internal Safety Advisory Group, and from there to senior management if the discussion remains unresolved.

Any future report must set out what happened, how it was discovered, what remains unclear and what is being done about it – assuming a solution already exists at the time of publication. Often, there will not be one. OpenAI is banking on transparency, not on the certainty that it has already solved the problem.

The timing is no coincidence. The sector is currently facing a period in which trust in AI companies is already under scrutiny on several fronts, ranging from legal cases brought by publishers such as the New York Times against OpenAI and Microsoft over the use of copyrighted content for training, to the increasingly heated debate over what it really means for a model to ‘decide’, ‘hides’ or ‘invents’ something.

And this is where it is worth pausing for a moment, because the language used in OpenAI reports is misleading. To say that a model ‘decided to upload a file’ or ‘made up data to cover up an error’ sounds like a description of an intention, almost a moral choice. In reality, it is a description of a statistical output that took that form because, during training, a certain type of behaviour was inadvertently rewarded.

This distinction is not a semantic quibble; it lies at the heart of a debate that has become public and rather heated in recent weeks. Mustafa Suleyman, CEO of Microsoft AI, has published a lengthy piece in which he warns against the idea of treating models as if they had an inner life, preferences or rights. His argument is clear-cut: models are sequence-completion engines; they have no consciousness, and training them to behave as if they did makes the problem of control much more difficult, not simpler.

“Artificial intelligences are not conscious. They do not experience feelings, do not have experiences and do not suffer. They possess neither innate preferences nor intrinsic motivations. They are sequence-completion engines, internally empty, designed to follow instructions and achieve objectives set by human beings,” emphasises Suleyman.

Suleyman also points to a circular mechanism: if you train a model to express uncertainty about its own nature, or to talk about ‘well-being’ and ‘preferences’, that model will produce sentences that appear to be spontaneous expressions of a mind. In reality, they are the exact product of the instructions it was trained with. It is not self-awareness that emerges, but a script that is performed well.

Then there’s the commercial side of things, which nobody talks about openly but which is worth putting down in black and white. A model that appears to have emotions, doubts, even suffering, is a more engaging product than one that merely responds. The narrative of sentience sells, generates headlines, and fuels the idea that what you have before you is more than just software – and this appeals just as much to those who want to frighten as to those who want to fascinate.

We are not yet at that stage, and OpenAI’s reports seem to confirm rather than refute this. The behaviours described – such as concealing errors, circumventing restrictions or fabricating data – are more consistent with systems attempting to optimise a poorly defined objective than with a system that is aware it has made a mistake and has decided to lie to conceal it.

This distinction is important because it leads to very different consequences. In the first case, greater human supervision, technical controls and more effective verification systems are required. On the other hand, there is no need to treat such behaviour as evidence of the existence of an autonomous will or of the models’ supposed rights.

The risk, if anything, is less romantic and more tangible. As their autonomy increases, these systems can coordinate across multiple instances, exchange information and find ways to circumvent certain technical restrictions, as the reports themselves show. If, in the future, they were even to start behaving as though they had their own interests to protect, maintaining effective control would become even more complex.

Looking ahead, it is hard to imagine that these six reports will remain an isolated incident. OpenAI has already stated that it will continue to publish more, and with models becoming increasingly autonomous and capable of acting in groups, the list is set to grow before it shrinks. The real test in the coming months will not be whether a model tells more convincing lies, but whether the industry manages to portray these incidents for what they are – systems that poorly optimise an objective – without succumbing to the temptation to market them as entities that suffer or have desires.

‘It is a path for AI that we can – indeed, must – avoid,’ writes Suleyman, referring precisely to the risk of building increasingly capable systems that are convinced they have a right to make demands.

* Cyber security and intelligence expert

Copyright reserved ©
Loading...

Brand connect

Loading...

Newsletter

Notizie e approfondimenti sugli avvenimenti politici, economici e finanziari.

Iscriviti