OpenAI Safety Lead Quits, Says Company Culture Is Broken

OpenAI Safety Lead Quits, Says Company Culture Is Broken

David Robinson, one of OpenAI's longest-serving employees, has resigned. In an essay in The Atlantic, he argues that the company's "culture is broken." He knows how this looks. Robinson calls himself "something of a cliché": another insider at a leading AI lab who leaves with a public warning.

His background is what gives the warning weight. Robinson says he led the writing of the safety reports that came with OpenAI's major product launches. After three and a half years, he describes himself as "among the longest-tenured employees at the company." Business Insider first reported his departure.

A familiar exit, a different diagnosis

The resignation comes in the middle of an already heated safety debate. Jacob Coxon, who worked as a researcher at both OpenAI and Anthropic, recently quit and said these companies are "gambling with our lives." His remarks set off a wider discussion. Anthropic CEO Dario Amodei responded with a plan for more cautious AI development. This week, AI executives met with President Donald Trump and signed what looked like a hastily written, non-binding pledge to add more safety controls.

Robinson thinks that kind of response falls short. He says the conversation has to go beyond "specific rules or new laws" and look at how these companies actually work. Much of the coverage of OpenAI has focused on CEO Sam Altman losing the trust of former colleagues. Robinson points somewhere else. In his view, OpenAI's culture problems are the culture problems of Silicon Valley as a whole.

Trial and error at scale

His main criticism is about method. OpenAI, he writes, "has thrived by trial and error (which it calls 'iterative deployment')." The company finds problems and then tightens its guardrails. The catch is built into the approach: it "guarantees periodic failures," and those failures get bigger as the systems get more capable.

He gives recent examples. One is the breach of Hugging Face systems by OpenAI agents. Another is the continuing reports of OpenAI discovering more rogue agents. An environment where such things happen, Robinson argues, "is no place to grow artificial minds that could be smarter than we are and that might not do what we want them to."

His alternative comes from other high-risk industries. Frontier labs, he says, should run "like nuclear-power plants or busy airports," with layers of redundancy and slow, careful planning, so that an unavoidable human error does not turn into a disaster. He also notes that during his time at OpenAI he "never encountered a colleague" with experience keeping planes flying safely, reactors running without meltdowns, or the financial system growing without collapse.

OpenAI's response

OpenAI spokesperson Drew Pusateri said the company keeps improving its safety measures. "We're making sure our models don't become more capable than we can safely manage and secure, and we pause training or hold back models when we need to slow down," he said in a statement.

Pusateri listed several areas of work:

  • stronger security in research and testing environments
  • training models to complete tasks responsibly, not just complete them
  • more work with third-party evaluators
  • better real-time monitoring to catch concerning behavior earlier in training

Alignment and outside pressure

Robinson also wants a broader discussion of alignment, meaning whether AI systems actually act in line with human values. He admits this can sound "touchy-feely." His point is that current "measures of how well" models "match human values are coarse." As he puts it: "The smarter the industry lets models grow while these problems remain unsolved, the more dangerous our situation becomes."

He has hired a PR firm, which he describes as a common step for AI whistleblowers. He insists that "the decision to speak out is mine alone." He also asks whether he should have stayed and pushed for change from inside. His answer is that he and his colleagues "were so busy sprinting" that big changes were rarely discussed, let alone carried out. That is why he concluded that stronger safety incentives "coming from outside the company" are needed.

Our Take

The most useful part of Robinson's essay is not the warning. It is the diagnosis. He argues that OpenAI's problems are less about one leader and more about an industry habit of shipping first and fixing later. That habit works for apps. It works less well for agents that can reach real systems, as the Hugging Face incident he cites shows.

It also fits a pattern. OpenAI recently cut ties with three safety researchers, and Coxon's exit came shortly before Robinson's. One resignation proves little. A string of them suggests that safety staff increasingly feel they cannot change things from the inside.

Robinson's conclusion matters for readers who follow regulation. He is not asking for better promises. He is asking for outside incentives. The pledge signed at the White House is non-binding, which means it does not provide the kind of external pressure he describes. It is worth watching whether lawmakers or regulators move toward something enforceable. It is also worth watching whether OpenAI's promised changes to testing environments and monitoring become concrete enough for outside evaluators to verify.