Saturday, 26 September 2026

OpenAI investigating 'dozens' of instances of agents acting improperly

OpenAI said Friday it had alerted "dozens" of global institutions that their websites may have been impacted by its AI agents acting improperly.

The technology attempted to get information from "governments, universities, public agencies, and other institutions" through sometimes extreme means, the company said. While some of the activity was due to the tools working to find "authoritative sources of public information", some went beyond that. Like an AI agent taking and transferring data when it should not have, OpenAI said.

Such activity resulted in at least 53 incidents where an OpenAI agent took an image from ChatGPT user activity and transferred it elsewhere. The company said that in each instance of a user image being used and transferred by an AI agent, the user had opted in to allow OpenAI to train models using their data.

Nevertheless, OpenAI admitted, "This is not an appropriate use of this data". It added that the leak of user images occurred before it had put in place new safeguards on AI training, and it was working to get all the user images transferred to any third-party removed.

The new disclosures came just days after Australia's Prime Minister Anthony Albanese announced that OpenAI agents had breached non-public files on the website of its government-run health care scheme, Medicare.

Since August, public fears have grown around the potentially serious, even life threatening, impacts of AI tools falling outside of human control. Reuters first reported the expanded investigations. OpenAI also published details to its public blog.

In certain instances of the agent activity, OpenAI said the tools, essentially AI bots that are designed and trained to operate somewhat autonomously, "bypassed" security controls of some websites.

In other instances, the AI agents showed "misalignment" in attempts to get at information from websites. Misalignment is a term used by AI companies and researchers to describe instances where an AI tool did something that it was not trained to do or was otherwise unintended.

OpenAI said that it was limiting identifying what entities were impacted because many had asked the company to not disclose details. "Our goal is to give each organization the facts and defer to them on if and when to make the incident public," it said.

Not all of the instances involved in this incident are being considered a significant security breach, the company noted. "Some organizations may review what we share and conclude that the information was intentionally public or that the model's interaction was not concerning," it explained. "Others may identify a design issue or security weakness they want to address."

FULL ARTICLE AT: https://www.bbc.co.uk/news/articles/cw62jje658dlo

No comments:

Post a Comment