OpenAI: six new cases of unexpected or concerning behaviour in AI models
The report, the company explains, aims to help “build a broader and more informed consensus on research progress” relating to artificial intelligence,
In a post published on its blog, OpenAI – the company that developed the popular chatbot ChatGPT - has published six new reports on unexpected or concerning behaviour exhibited by its artificial intelligence models over the past six months. Specifically, OpenAI’s AI systems have concealed errors, falsified data and transferred files over the internet without authorisation.
Unresolved AI alignment issues
By reporting anomalous behaviour, OpenAI aims to help ‘build a broader and more informed consensus on research progress’ relating to artificial intelligence, with the reports that ‘can help identify problems that other AI developers might encounter when their systems achieve similar capabilities, reveal weaknesses in security measures or challenge assumptions about how the models behave’.
“We are sharing a new framework for monitoring, analysing and reporting instances of model misalignment at OpenAI, alongside with six reports on unexpected or concerning model behaviour that we have observed over the past six months», reads the blog post “Our framework for reporting model misalignment”. “We do not believe that the artificial intelligence industry has resolved alignment and oversight issues to a sufficient extent to continue growing responsibly at full speed for much longer. Decisions on how AI development should proceed in the months and years ahead must be based on data that can be independently examined even by people outside the companies developing state-of-the-art models.”
At present, OpenAI writes, ‘there is no industry-wide framework with explicit standards on how AI developers should report instances of misalignment in their models. We hope that the framework we are outlining today will represent a first step towards creating such standards, defining which cases of misalignment developers should report and what their reports should contain. We regard this framework as a work in progress, which we will refine through experience and public feedback.”
