
TLDR: OpenAI published on August 26, 2025 that, in cases of an imminent threat of serious physical harm to other people, conversations may be reviewed by humans and, as a last resort, forwarded to authorities. Cases of self-harm are not forwarded to the police, according to the company.
Recently, an article came out on Futurism about a new function in ChatGPT. Yes, your conversations can now be “scanned” and reported to the proper authorities.
Considering the recent controversies involving OpenAI, from user psychosis to more serious cases such as the death of a teenager, it makes sense that this change is already in effect.
OpenAI created a post titled, “helping people where they need it most.” In the post, OpenAI says, literally:
- “When we detect users planning to harm other people, we route the conversation to specialized pipelines reviewed by a small team… If the reviewers determine an imminent threat of serious physical harm, we may forward it to authorities.”
- OpenAI also states that it is not forwarding cases of self-harm to the police “to respect privacy,” and that the model tries to redirect people showing suicidal ideation to help resources.
It is worth remembering that this was already in OpenAI’s privacy policy. The idea of preventing a greater danger when some user tries to carry out or asks about something genuinely dangerous seems quite interesting, but some things were not well defined, such as:
- Objective criteria: what exactly qualifies as an “imminent threat”?
- Scope limits: how far does “dangerous content” go?
- False positives: can irony, role-play, or artistic venting be misinterpreted?
- Transparency: are there auditable logs, user notification, and an appeal process?
On top of that, if your data can be used indirectly when there is an imminent threat, what guarantees that eventually they cannot be used at other times? Could an order from the American government also change what is considered a threat?
Of course, it is important to have parental controls and obviously a system to try to prevent possible threats, but is there not a lack of more transparency and auditing to really define what does or does not go to the authorities? Not to mention that this literally goes against the idea of privacy that OpenAI had been talking about from the beginning.
We will possibly have news about security/control systems in the not-so-distant future at OpenAI, Google, Anthropic… But it is necessary to be careful so that something that could really be positive, like preventing a catastrophe, does not turn into an excuse to violate users’ privacy.
An “emergency brake,” if used the wrong way, can bring more problems than solutions.