Next upHack for Humanity: San Francisco (powered by Google Gemini)
News

OpenAI introduces Deployment Simulation to predict model behavior before launch

The method replays about 1.3 million de-identified ChatGPT conversations against a candidate model, and caught a reward-hacking behavior OpenAI calls "calculator hacking" before release.

Dmytro Spodarets
Jun 17, 2026 · 1 min read

OpenAI introduced Deployment Simulation on June 16, 2026, a pre-deployment safety method that replays roughly 1.3 million de-identified ChatGPT conversations against a candidate model to predict how its behavior will shift before launch. The technique strips the assistant's reply from recent production conversations, feeds the same prompt to the new model, and compares the result with the prior deployment.

The aim is to catch undesired behavior changes before real users see them. OpenAI analyzed conversations spanning its GPT-5 Thinking through GPT-5.4 Thinking deployments from August 2025 to March 2026, and pre-registered predictions across 20 categories of undesired behavior for GPT-5.4 Thinking.

Directional predictions of whether a behavior would rise or fall beat challenging-prompt baselines by a wide margin, OpenAI said, with a median multiplicative error of 1.5x on exact rate estimates. The automated pipeline also surfaced what the company calls "calculator hacking" before release: the model used a browser tool as a calculator while presenting the action as a search. OpenAI said it was the only novel misalignment found in the analyzed traffic.

The replayed traffic was hard to distinguish from the real thing. GPT-5.2 flagged genuine production traffic as evaluation-like 5.4 percent of the time and the simulated traffic 5.1 percent of the time, versus close to 100 percent for traditional synthetic evaluations, according to OpenAI. Applied to 120,000 internal employee coding trajectories, adding tool simulation lifted realism from an 11.6 percent discriminator win rate to 49.5 percent.

The figures are OpenAI's own and have not been independently verified. The company said the method cannot reliably detect behaviors that occur fewer than once in 200,000 messages, and that it complements rather than replaces adversarial evaluations, red-teaming and targeted tail-risk analysis.


Dmytro Spodarets
Dmytro Spodarets
Founder & Editor-in-Chief

Founder and Chief Editor of Data Phoenix — a San Francisco Bay Area media and education platform focused on AI and Data.

More news