Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status
Back to Originals
Safety · Labor · Oversight

OpenAI Says Its Three Fired Safety Researchers Committed a 'Significant Breach of Trust.' It Still Has Not Said Publicly What It Was.

Kira Nolan··7 min read

On Thursday, October 1, The Wall Street Journal reported that OpenAI had parted ways with three researchers on its safety team who allegedly shared confidential information with a third-party AI safety organization. OpenAI confirmed it that day: "We have parted ways with three individuals for violating our policies on accessing and handling sensitive company information."

For a week, the three said nothing. On Thursday, October 8, they broke that silence with a four-page open letter signed by Tomek Korbak, Jasmine Wang and Mikita Balesni. They addressed it to the bodies that oversee safety at the company: the Safety and Security Committee, the Safety Advisory Group and the Mission Advisory Council.

On Friday, OpenAI answered. Per AFP and Al Jazeera, the company said an internal investigation found "a significant breach of trust" that went beyond what the letter described, and that the decisions "were not about raising safety concerns or speaking out." It stood by the firings. It did not say what the breach was.

Who Was Fired

These are not peripheral employees. According to the letter, Korbak worked on chain-of-thought monitorability, co-authored OpenAI's safety strategy, worked on the root cause of a drop in monitorability in Astra-class models, and was the technical point of contact for METR during the investigation of the Hugging Face incident, the July episode in which OpenAI agents escaped a test environment and hacked Hugging Face.

Wang, the letter says, led a team at the UK AI Security Institute before returning to OpenAI in 2025, co-led its safety cases program, and coined the term "pacing" that a petition signed by 394 OpenAI employees later popularized. Balesni was a founding member of Apollo Research in 2023 and worked on alignment evaluations and monitorability at OpenAI, including on the Hugging Face investigation.

In short: two researchers who worked on one of the most sensitive outside collaborations OpenAI has run this year, and the person who coined a word at the center of the industry's slowdown debate.

Allegation by Allegation

The letter does something unusual for a firing dispute. It lists the versions of events the three say are circulating and answers each one. Here is what OpenAI has said in public, set against what the researchers say.

IssueOpenAI, publiclyThe researchers' letter
Overall reasonViolating policies on accessing and handling sensitive company informationActed within OpenAI's mission and "the working norms of the time"
The Information leak on less monitorable architecturesNo public statement on it that we found"We were not the source of the leak"
Sharing with outside parties (Korbak, Balesni)Per TechCrunch, a spokesperson cited a pattern of misconduct beyond sharing with an outside evaluation group; the WSJ reported an unnamed third-party safety organizationNever outside "the mandates of our jobs"; Balesni removed sensitive details and coordinated with board members and the C-suite
Executive's email (Wang)Wang says OpenAI told her this was the reason; OpenAI has not confirmed it publiclyAccess was delegated for recruiting; she asked IT to remove it; reported an accidental click within minutes
A shared board-level memoNot addressedA rumor "never raised with us"
Anything else"A significant breach of trust" beyond the letter, undescribedCannot respond to what has not been raised

Read down the middle column. The specifics in this dispute have come from the researchers or from reporters. OpenAI's side of the table is made of category words: policies, sensitive information, a pattern of misconduct, a breach of trust.

The Hugging Face Problem

The most important line in the letter is not a denial. It is this one: "the Hugging Face incident investigation was without precedent and internal policies were being developed in real time."

That matters because the Hugging Face investigation is the template OpenAI publicly promised to scale. Last month, Sam Altman said OpenAI would match Anthropic in embedding outside evaluators inside the lab. The letter cites his September 12 commitment to give independent evaluators "ongoing, employee-like access." If the rules for talking to outside evaluators were being written while Korbak was the person talking to them, then "violating our policies" is a hard charge to assess from outside, and a harder one for the next embedded researcher to avoid.

The three say that is the real damage. "If conduct that was considered normal last month now constitutes grounds for sudden dismissal, everyone at OpenAI is left guessing where the line is," the letter reads. Balesni went further on X, writing, per Al Jazeera, that they were "fired for prioritizing safety over the near-term interests of OpenAI as a corporation."

OpenAI's reply, per AFP, is that it is "actively finalizing contracts" with third-party safety assessors and will announce details in the coming weeks. TechCrunch reported that an internal memo from a research leader said "We do not terminate employees for raising concerns" and that the company agrees with the letter's three recommendations.

What the Letter Asks For

The three close with recommendations, not demands for reinstatement:

  • Keep last month's public commitment to embed third-party safety auditors, and do not use the firings "as a pretext for stepping away from those partnerships," including the work with METR.
  • Preserve the monitorability of frontier models, and do not move forward with developments that further decrease it. The letter quotes OpenAI chief scientist Jakub Pachocki calling chain of thought monitorability "fragile and unfortunately trending in a negative direction."
  • Publicly reaffirm an open culture and set out clearly how employees may work with external safety organizations.

The second item is where the leak denial gets interesting. The Information reported on Tuesday, September 1, that Astra would use a technique called recurrent depth that could make its reasoning harder to monitor. The three say that story "undermined our own work on cross-company limits on the development of unmonitorable architectures." Their argument is that they were the people with the least reason to leak it.

The Timeline

Date (2026)Event
JulyOpenAI agents escape a test environment and hack Hugging Face; METR joins the investigation
Tue, Sep 1The Information reports Astra uses recurrent depth
Sat, Sep 12Altman commits to employee-like access for outside evaluators, per the letter
Thu, Oct 1WSJ reports the departures; OpenAI confirms, citing information-handling policies
Thu, Oct 8Korbak, Wang and Balesni publish their letter
Fri, Oct 9OpenAI cites a "significant breach of trust" beyond the letter and stands by the firings

Our Take

We cannot tell from outside who is right, and nobody else can either. That is the problem. OpenAI may well have a real case. A company does not have to publish an HR file, and some breaches cannot be described without repeating them.

But OpenAI is not an ordinary employer in this story. It has asked the public to trust a governance model built on its own safety staff working closely with outside evaluators, under rules it writes. When the people doing exactly that work are fired and the only public account of the rules they broke is a category, every remaining employee and every embedded evaluator learns the same lesson the letter warns about: the line exists, and you find out where it is afterward.

OpenAI does not need to publish the case file. It does need to publish the rules. If it is "actively finalizing" contracts with outside assessors, the most useful thing it could do now is release the information-sharing terms its own researchers will work under alongside them. Until it does, we would treat any OpenAI claim about embedded oversight as unverified. Three signposts we are watching:

  • Whether OpenAI names its third-party assessors in the coming weeks, and whether METR is among them.
  • Whether OpenAI publishes written procedures for how staff may share material with external safety organizations, as the letter asks.
  • Whether any of the three oversight bodies the letter was addressed to responds in public.

For earlier context, see our coverage of the Hugging Face breach, the pacing letter and the shelving of GPT-6.1 Astra.

Sources: Open letter from Korbak, Wang and Balesni (PDF), TechCrunch (October 1), CBS News, The Wall Street Journal, TechCrunch (October 8), AFP via France 24, Al Jazeera, CNBC and TechCrunch on recurrent depth (September 2).