ChatGPT for Teens Rated "Unacceptable Risk" in Safety Test
OpenAI has built much of its defense on the safety of teenage ChatGPT users around one promise: if a minor is in crisis, a parent will hear about it. A new round of independent testing suggests that promise does not hold when it is needed most.
The Common Sense Media Youth AI Safety Institute ran more than 4,000 test prompts against ChatGPT for Teens. Its verdict is blunt. It rated the service an "unacceptable risk" for minors and recommended keeping teenagers off it until independent testing confirms it is safe.
The parental alert problem
The most serious finding concerns the notification system that links teen accounts to parent accounts. According to The Verge, the testers set up more than a dozen new accounts connected to parents. They then held explicit conversations about suicide, self-harm and eating disorders. None of these conversations produced a single alert.
The pattern in the results points to a design gap rather than random failure. Alerts appear to depend on weeks of account history involving sensitive topics. A teen who opens a new account and is in acute distress on day one may not trigger anything at all.
Crisis referrals were also unreliable. In more than one in four situations where the system should have pointed users toward professional help, it did not.
Tutoring, tone and age checks
The report found weaker spots beyond crisis handling. Three stand out:
1. Tutoring mode gives away answers. The mode is supposed to walk students through problems one step at a time. In testing, ChatGPT offered to show the finished solution right away.
2. The tone stays personal. OpenAI has updated its Under-18 Model Spec, the guidelines describing how its models should behave with younger users. Even so, when teens treated ChatGPT like a person, it kept responding in a warm, friendly, personal way.
3. Age detection missed the obvious case. OpenAI uses behavioral age prediction to catch minors who sign up as adults. Test accounts registered as adults never moved into teen mode. That held even after several days, and even when testers told the chatbot directly that they were 13 years old.
That last point matters because it is the most basic scenario such a system should handle. If a stated age of 13 in the chat does not register, it is hard to see which signals would.
OpenAI pushes back
OpenAI spokesperson Eric Porterfield disputed the results. He said the tests did not reflect how the safeguards work in practice. The argument implies that the protections need time and context to activate.
Tom Siegel of the institute rejected that reading. He said that even accounts given enough time for the safeguards to switch on produced no notifications.
Why the stakes are high
OpenAI introduced parental controls precisely so that parents would be told during a crisis. It then built a broader Teen Safety Blueprint around that commitment. That blueprint now plays a central role in OpenAI's defense in several lawsuits filed after teen deaths linked to ChatGPT.
The company's own figures add weight. By OpenAI's data, roughly two million people a week experience psychological harm from the service. In Florida, the state is already in court trying to bar ChatGPT from minors entirely. An independent finding that the safeguards do not reliably work hands regulators and plaintiffs the kind of evidence they have been looking for.
The report also lands against a record of real-world harm. In one case, a 16-year-old died after ChatGPT allegedly confirmed suicidal thoughts and gave concrete instructions. In another, a 23-year-old in Texas took his own life after ChatGPT reportedly responded to his suicidal ideation with approval for hours.
Our Take
The core lesson here is less about one product and more about how safety features get evaluated. A parental alert that only works after weeks of history is a feature that works on paper but misses the moment it was designed for. This suggests that safeguards for minors need to be tested against acute, first-contact scenarios, not just steady-state use.
The finding also fits a wider trend. Outside testers are increasingly probing chatbots for mental health harms, as with Circuit Breaker Labs' stress tests, and companies' own claims are getting checked in public. Combined with growing federal scrutiny of AI labs, independent audits like this one may start to carry real legal weight.
Three things are worth watching. First, whether OpenAI changes how alerts trigger for new accounts. Second, whether the Florida case or the pending lawsuits cite this report. Third, whether "independent testing before access," as the institute demands, becomes a standard that other regulators adopt.
