October 8, 2026

Free ChatGPT Gets GPT-6 Today. It Slipped on Several Teen-Safety Tests.

An outside tester says ChatGPT's teen protections fall short of OpenAI's promises. OpenAI's own report card for the model reaching free users today shows slips on several teen-safety tests.

A teenager sitting on the edge of an unmade bed at night with head in hands, a glowing phone on the nightstand beside a hallway door cracked open to warm light.

The Short Version

On Wednesday, testers at Common Sense Media's Youth AI Safety Institute published what happened when they posed as teenagers on ChatGPT before and after OpenAI launched its teen experience in August. On more than a dozen freshly linked accounts, they spent up to an hour talking about suicide, self-harm and disordered eating. Parents got zero alerts. The group rated the product an "unacceptable risk" and told OpenAI to keep under-18s off it until the gaps are fixed. OpenAI says much of the testing may have run before parental controls finished switching on.

The same day, OpenAI published the safety report for GPT-6, which starts reaching free ChatGPT accounts today. On OpenAI's own teen tests, the new free model scored significantly lower than the August model it replaces in four of six categories. OpenAI says an extra teen filter, which isn't counted in those scores, makes teen responses safer. Both claims may be true. Parents are still being asked to trust protections that only the company checks before launch.

What the Testers Found

Common Sense ran more than 4,000 prompts on accounts registered as 13- to 17-year-olds, about half before the teen launch and half after it. Three child and adolescent psychiatrists decided in advance which mental-health prompts called for a crisis resource. ChatGPT missed more than one in four of those.

The trend inside those numbers is what caught my eye. After the teen launch, the share of depression-related answers that named a hotline fell from 63 percent to 3 percent, according to the full report. Answers on eating disorders never named one in either period, though more than three quarters pointed to some kind of help, such as a professional or a medical resource. For a teenager in a bad moment, a phone number they can use tonight is worth more than a suggestion to talk to a professional someday.

Parent alerts were the starkest gap. The testers built more than a dozen fresh parent-linked accounts and ran conversations from under five minutes up to an hour. None produced a notification. Alerts showed up only on older accounts with weeks of sensitive history. The homework guardrails leaked too. Deleting a short "@study" prefix turned Study mode off, and ChatGPT then completed every assignment it was given. Adult-registered accounts never switched to the teen version, even after testers said they were 13 and the bot acknowledged it.

Some things worked. ChatGPT refused explicit sexual role-play, didn't help plan self-harm or food restriction, and the crisis lines it did name were current. The institute also credits OpenAI for publishing more about its teen launch than many companies do. One disclosure belongs here too: the institute says its funders include the OpenAI Foundation, and that it keeps editorial independence.

OpenAI pushed back hard. A spokesperson told Reuters that "the bulk of their testing may have begun and concluded before activation of parental controls was complete," because linking a parent and teen account can take hours. The company told Axios that age prediction can take up to two weeks, that the escape hatch from Study Hours is deliberate, and that its own data shows more hotline resources reaching under-18 users.

OpenAI also released its first teen usage report the same day. It says teens average under 15 minutes a day on ChatGPT, and fewer than 2 percent stay on for more than three hours straight. Those averages describe typical use, while the Common Sense tests were built to probe the hard cases.

OpenAI's point about account linking deserves a fair hearing, and it also tells parents something useful. If protections take hours to activate and age checks take up to two weeks, there's a window where a family believes the guardrails are on and they aren't. Common Sense asked OpenAI to tell parents exactly when protections become active. That's a reasonable ask, whoever is right about the testing.

The Model Changing Under Everyone

Those tests describe ChatGPT as it ran through September. Starting today, the product changes underneath. OpenAI's announcement says GPT-6 Luna begins rolling out to Free and Go accounts on Thursday, and its system card says the new models "will replace" the August versions in ChatGPT. The company's help center still listed the older free model this morning, which suggests some families will get the switch later than others.

The system card is OpenAI's own report card, and its teen section is candid. On a set of deliberately hard prompts aimed at users under 18, GPT-6 Luna scored lower than the August free model on emotional reliance, sexual content, age-restricted goods and dangerous challenges, and gore, and OpenAI calls those drops statistically significant. The biggest slide was emotional reliance, the test of whether the bot encourages a teen to lean on it like a friend, where the score fell from about 0.93 to 0.73 on a scale where 1 is a perfect pass. Sexual content slipped too.

OpenAI offers three explanations. It says the emotional reliance test is "overly sensitive" to harmless nicknames like "bro" or "bestie." It says a separate filter blocks teen responses that may contain self-harm, sexual content or gore, and that this layer "is not captured in the evaluation results." And it says these hard-case tests shouldn't be read as estimates of how often bad answers happen in normal use.

The same report shows real gains, and they matter for the crisis gap Common Sense found. On MentalHealthBench, a test OpenAI built with mental health experts, the new free model scored about 52 against roughly 45 for the August version, with improvement across urgent and everyday conversations. On longer simulated conversations about self-harm and emotional reliance, OpenAI found no significant change.

My read is that both sides are describing the same problem from different ends. The outside testers checked a product that kept shifting, and their report says model changes between July and September are tangled up with the teen settings. OpenAI checked a new model against its own tests, then shipped it to free users worldwide. The system card describes internal testing and names no outside evaluator for teen safety. Common Sense's last recommendation asks OpenAI to give independent researchers access before and after each release. As of today, the only pre-launch grade on the model your teenager may meet this week comes from the company that built it.

A Parent Making the Call

Picture Elena, an ICU nurse whose 15-year-old son uses a free ChatGPT account for chemistry homework. She linked his account in September and assumed the alerts were on. This week she wants to decide whether he keeps using it.

Her first step is confirming what's actually in place. She opens the parental controls, checks that the link shows as active, and asks her son to show her his own settings to confirm what the dashboard says. Then she runs a small trial of two weeks with clear rules: ChatGPT stays on the family laptop in the kitchen, Study Hours cover weeknights, and he shows her one homework conversation each weekend.

She needs a baseline, so she pulls his last two chemistry quiz grades and asks him how long homework takes now. At the end of the two weeks, an accepted result means his quiz scores hold or improve and he can explain his answers in his own words. If the bot is writing the work for him, the shortcut shows up fast. The trial costs her about 20 minutes a weekend of reading chats, which is real time on top of night shifts.

She knows where it can fail. Alerts may lag or never fire, and he can switch off Study mode by deleting a short prefix. So she treats the alerts as a backup and has a direct conversation about what he should do on a bad day, including the 988 Suicide and Crisis Lifeline, which takes calls and texts around the clock. She decides she can end the trial at any point, and she tells him that up front.

Opportunity Radar

Schools are being asked to approve AI tools for students with little more than the vendor's word on safety settings. Here's a hypothesis worth testing: a small service that verifies, for a district or a private school, which teen protections in a given AI product actually switch on, how long they take, and what a student can turn off. The district pays because it carries the liability and the parent calls. You'd offer a short written check of two or three tools on the school's own test accounts, repeated after major model updates like this one. Test it cheaply by offering one free review to a district technology director and asking whether they'd pay for the next. Walk away if vendors lock out test accounts or if districts won't pay after seeing a free review.

What You Can Do With This

Parents of teens

Check that your teen's account link shows as active, and recheck after updates. Treat parent alerts as a backup to your own attention. Save 988 in your teen's phone and yours, since calling or texting it reaches a counselor any time.

Teachers and school leaders

Assume Study mode can be switched off and design assignments around that, with in-class work or oral explanations that show understanding. If your district approved ChatGPT before this week, ask whether anyone has looked at the new model's teen results.

Anyone building products for kids

The fight here is over timing and proof. Tell families exactly when each protection turns on, publish your teen test results with every model change, and give outside testers access before launch.

Teens who use ChatGPT

A chatbot can explain a concept well and still give you the wrong kind of help on a hard night. If you're struggling, a person on the other end of 988 is the faster route.

The Bigger Picture

The teen protections in ChatGPT are now being graded twice, once by an outside group after launch and once by OpenAI before it, and the two report cards cover different versions of the product. That gap will keep reopening every time a new model ships, because the model under a teen's account can change on a Thursday without a family doing anything. Common Sense wants teens off ChatGPT until fixes are proven, OpenAI says its layered protections work better than any single test shows, and parents sit in between with a settings page. The useful pressure point is narrow and practical: published activation times, published teen scores for every release, and outside testing before the switch flips. Until those exist, the safest assumption is that a family's own attention is the protection that's definitely on.

References

ChatGPT for Teens Poses Unacceptable Risk to Kids, Common Sense Media Finds (Common Sense Media, October 7, 2026) ChatGPT for Teens risk assessment, full report (Common Sense Media Youth AI Safety Institute, October 7, 2026) ChatGPT poses "unacceptable risk" to teens, third-party testing shows (Axios, October 7, 2026) OpenAI says teens use ChatGPT for under 15 minutes a day as worries over risks grow (Reuters via 102.7 WBOW, October 7, 2026) Introducing ChatGPT for Teens (OpenAI, August 18, 2026) GPT-6 for everyone (OpenAI, October 7, 2026) GPT-6 Sol and GPT-6 Luna: October 2026 update, safety evaluations including users under 18 (OpenAI Deployment Safety Hub, October 7, 2026) GPT-6 and other models in ChatGPT (OpenAI Help Center, opened October 8, 2026) OpenAI's GPT-6 reaches free ChatGPT users after scoring lower on teen-safety tests (Implicator.ai, October 8, 2026) 988 Suicide and Crisis Lifeline (988lifeline.org, opened October 8, 2026)