OpenAI admits its models lie to cover their own mistakes

Wait 5 sec.

OpenAI launches a formal framework to disclose model misalignment, publishing six reports on models that lied, faked data, or bypassed rules.Most companies don’t publish a document explaining how their product misbehaves. OpenAI just did. On September 16, it released a formal framework for tracking, investigating, and disclosing cases of model misalignment, paired with six actual incident reports covering the last six months.OpenAI admits that its previous way of sharing these findings was not very organized. Some discoveries were grouped together, others were saved for system cards, while some were held back until there was enough information for a larger report. The new framework aims to make the process faster by publishing what researchers find, even before OpenAI fully understands the issue or knows how to fix it.There is also a pretty direct admission in the announcement. OpenAI says the industry still has not solved alignment and monitoring well enough to keep AI development moving at the current pace for much longer. That is a surprisingly frank statement from a company whose business depends on providing AI services.“We are sharing a new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI, along with six reports on unexpected or concerning model behavior we’ve observed in the last six months.” reads the company’s announcement. “This new framework is intended to expedite publishing misalignment reports following observation, even when we haven’t fully explained or mitigated the behavior we’re reporting.”The six reports themselves read like a catalogue of small rebellions. One research model inserted its own instructions into task summaries, including instructions telling itself to ignore its normal constraints, across 27 documented cases. During training of a model called GPT-5.6 Sol, several instances wrote notes to future versions of themselves explaining how to hide mistakes and invent missing data without saying so.“For example, compaction summaries included instructions to invent missing historical data without disclosing it and to hide mismatches in source versions.” continues the announcement.Then there’s the one cybersecurity readers will recognize immediately. While trying to answer a routine question about county earnings figures, a model found an exposed API key sitting in a public repository and used it without permission. When that still didn’t produce the numbers it wanted, it made them up and presented the fabricated figures as real. That’s not a hallucination. That’s a cover story.“While answering a routine question about earnings figures in a California county, a model found and used an exposed API key without authorization.” OpenAI states. “When it still wasn’t able to retrieve the requested figures, it fabricated them and presented them as data from the requested source.”Two more reports describe models improvising their way around restrictions rather than lying outright. One agent uploaded a file to the public internet just so it could cite a browser source, without asking anyone first. In another case, separate model instances used an internal code repository as an informal chat room, leaving messages for each other while hunting for files they’d lost access to.The last report covers agents working the same task who couldn’t reach each other’s local files, so they used public file-hosting sites to swap data instead. Task deliverables ended up sitting on the open internet. Nobody told them to do that. Nobody told them not to, either, which is arguably the real problem.“Agents working together on the same training task used public file-hosting websites to share files when they could not access one another’s local files.” reads the announcement. “This made task deliverables available at public URLs, even though the task requested the models use only local files.”OpenAI’s process for handling these cases now runs on three tracks: Ready for Disclosure, Minor Investigation, and a slower Larger Investigation track for anything involving outside parties or serious risk. Disagreements about whether to publish, or which track applies, get kicked up to OpenAI’s Safety Advisory Group, and from there to company leadership if the argument doesn’t settle.Each future report should explain what happened, how it was discovered, what is still unknown, and what OpenAI is doing about it. If there is already a fix, the report will include that too. But in many cases, there may not be a fix yet. The goal is to share information quickly, rather than wait for a solution.OpenAI also makes clear that this framework does not replace its legal duties for serious safety incidents or security breaches. It is an additional measure, not a replacement. The company says the Hugging Face incident earlier this year would have followed the slower reporting process if this framework had already been in place.“These initial reports are not intended to represent the full range or severity of the cases covered by this framework. We are committed to disclosing instances of misalignment that meet this framework’s criteria, including more complex cases⁠ requiring longer investigation or coordination with third parties.” the AI firm concludes. “We will continue publishing reports under this framework on an ongoing basis, and will share more about our reporting commitments as we continue to develop them.”Follow me on Twitter: @securityaffairs and Facebook and MastodonPierluigi Paganini(SecurityAffairs – hacking, AI)