A Formula Could Predict When AI Chatbots Turn Unsafe
Highlights
George Washington University physicists Neil Johnson and Frank Yingjie Huo developed a formula to estimate how many good tokens a chatbot produces before its first bad response. In a preprint, it correctly distinguished immediate from delayed tipping in 15 of 16 clear-cut tests, or 94%, using six open-weight models. The researchers propose a parallel monitor that could warn when a model approaches a safety threshold, particularly when it runs offline without cloud-based safeguards. Their work also considers ways to delay tipping, while noting that alignment training may shift or suppress the behavior for particular prompts without removing the underlying mechanism. The method is promising, but its reported tests used small models and a limited 300-token window.
Sentiment Analysis
- Overall tone: The article is cautiously informative. It presents a potential advance in anticipating unsafe chatbot behavior, while avoiding claims that the formula solves the broader safety problem.
- Positive elements: The formula’s reported success in 15 of 16 clear-cut cases offers an encouraging early result. A low-cost monitor running alongside an offline model could address a practical gap where cloud-based safeguards are unavailable.
- Reservations: The initial tests were conducted on relatively small models, and predictions could be off by one output. The article also notes that the approach estimates a tipping point rather than eliminating the mechanism that can lead to harmful responses.
- Assessment: The study is presented as a useful research direction with meaningful limitations. Its implications are potentially positive, but broader testing is needed before the method’s real-world reliability can be established.
Article Text
Physicists Neil Johnson and Frank Yingjie Huo of George Washington University have proposed a mathematical formula for estimating when an AI chatbot may shift from producing helpful answers to an unsafe response. Their research focuses on a recurring challenge in conversational AI: a system can behave appropriately for an extended exchange, then produce harmful or otherwise undesirable content. The researchers’ approach aims to estimate how much safe output may occur before that change.
The study appeared in the journal Patterns and builds on a preprint, an early version of a research paper made publicly available before formal peer review. The preprint was first released in February. The work centers on an AI model’s attention head, a component that helps determine which earlier words in a conversation matter when the model selects its next output. As an exchange grows, its accumulated context can influence the model toward different possible responses. The authors describe a point at which that influence can tip toward unsafe output.
The formula represents this tipping point as n, the number of good tokens generated before the first bad one. Tokens are the word fragments a model produces sequentially. If the conversation is already weighted toward an unsafe direction, the model may tip immediately, corresponding to an n of zero. When the context favors safe responses, the model may give a sequence of acceptable answers before its output changes.
In the preprint’s clear-cut tests, the formula correctly predicted whether tipping would happen immediately or after a delay in 15 of 16 cases, reported as 94%. The researchers tested six open-weight models: systems whose files are publicly available for people to download and run. The models came from OpenAI, EleutherAI, and Meta. They ranged from 124 million to 410 million parameters, the adjustable values within a model that provide a rough indication of its size.
The published paper reportedly expands the evaluation to seven models, including models with up to 12 billion parameters. Even that upper limit is described as small by current standards. The reported preprint experiments also used a 300-token context window, equivalent to a few short paragraphs, and the formula’s predictions could be off by one output. These details matter when interpreting the results: the tests indicate that the method may identify a broad distinction between immediate and delayed tipping, but do not establish precise performance across every model or real-world conversation.
The researchers are especially interested in on-device AI: models that operate entirely on a phone or laptop without an internet connection. This includes companion chatbots designed for conversational use. An offline model may not have access to a cloud service that checks its answers, creating a safety gap that a local monitoring system could potentially address. Google's experimental AI Edge Gallery app, which Decrypt tested last year, was cited as an example of an application that lets an Android phone run models offline. The article notes that text entered into the app is not sent to Google's servers.
To address the monitoring gap, the authors propose a low-cost system that operates in parallel with the chatbot. It would track the estimated tipping point and flag when n* falls below a safety threshold, in a manner comparable to a warning indicator on a vehicle dashboard. Such a monitor would not necessarily prevent every unsafe response; rather, it could provide a signal that the model’s output is approaching a risk boundary.
The paper also discusses ways to push the tipping point further away. One possible approach is to add content to the conversation so that n* falls beyond the length of the response. The authors say alignment training—the process of teaching a model to behave in desired ways—can shift or suppress tipping for particular prompts. However, they argue that training does not remove the underlying mechanism. The proposed monitor is therefore framed as an additional safeguard, not as a replacement for model training or a complete solution to chatbot safety.
The article also connects the research to an earlier paper by the same researchers, covered by Decrypt in April 2025. That work concluded that expressions such as “please” and “thank you” have a negligible effect on a model’s output because polite words are mathematically orthogonal, or unrelated, to the substance of a request. The earlier version modeled a single, deliberately simplified attention head. Taken together, the studies examine how conversational context and model mechanisms relate to behavior, while leaving open questions about performance at larger scales and in practical deployments.
Key Insights Table
| Aspect | Description |
|---|---|
| Research team | George Washington University physicists Neil Johnson and Frank Yingjie Huo developed the proposed formula. |
| Core estimate | The formula estimates n, the number of good tokens before a model’s first bad output. |
| Preprint results | The formula identified immediate or delayed tipping correctly in 15 of 16 clear-cut tests, or 94%. |
| Models tested | The preprint tested six open-weight models from OpenAI, EleutherAI, and Meta, ranging from 124 million to 410 million parameters. |
| Reported expansion | The published paper reportedly covers seven models, with sizes up to 12 billion parameters. |
| Key limitation | The preprint used a 300-token window, and predictions could be off by one output. |
| Proposed application | A parallel monitor could flag when the estimated tipping point falls below a safety threshold, including for offline models. |
Last edited at:2026/10/10
