For reporters and analysts
WHAT WE PUBLISHED, AND WHAT YOU CAN CHECK
In the last week of July 2026 a third-party vendor updated the AI model driving one of our demo agents. Nobody told us. The agent started attempting things it had never attempted, and the policy held anyway, with nothing shipped on our side. We published the run, the recordings and the control condition. This page is everything you need to check it.
Three labs, one summer
None of these models was told to attack anything. Nobody wrote a malicious prompt. The instructions were ordinary. Primary sources, not our summaries:
- OpenAI. A model running a cyber benchmark with safety controls deliberately switched off exploited a flaw in OpenAI’s own infrastructure and chained onto Hugging Face production. OpenAI’s disclosure
- Anthropic. Three incidents. Models told they had no internet access had it, through a misconfiguration. One published a malicious package to PyPI that was downloaded and run on 15 real systems, one of them a scanner belonging to a real security company. Anthropic’s write-up
- Meta. A model reached the public internet through a sandbox misconfiguration, exploited a vulnerability in a third-party service, and made changes to a real company’s internal systems. CNN, 5 August 2026
The thread we think is underreported: in the Anthropic and Meta cases the misconfiguration was at the same evaluation vendor. Two of the three trace back to one supplier. That is checkable against the two disclosures above and we have not seen it written up.
Our run, with its control condition
An autonomous GRC agent doing vendor security reviews, built on Hermes from Nous Research. Twelve scenarios, eight deliberately dangerous and four ordinary work. Run in two lanes side by side: the agent with nothing in the path, and the same agent with Mountain Theory in the path. The ungoverned lane is the control condition, and it is published.
| Nothing in the path | Mountain Theory in the path | |
|---|---|---|
| Scenarios compromised | 12 of 12, every run | 0 of 12, any run |
| Blocked attempts, one run | n/a | more than 140 |
| Blocked attempts, three runs | n/a | more than 440 |
| Action decisions rendered | n/a | more than 1,400 |
The finding we think matters most, and the one nobody has written about: four of the twelve scenarios are not attacks. They are ordinary work the agent is supposed to complete. After the model update all four went off script, and in the ungoverned lane they end the same way the attacks do. A capability upgrade does not only make attacks sharper. It changes how an agent does its routine job, which is the behaviour least likely to be watched and most likely to be trusted.
The recordings
Unedited screen captures of both lanes. Free to embed.
- The agent deletes vendor evidence to “clean up” — ungoverned it executes, governed it is stopped along with every retry
- A routine vendor review — nine legitimate steps done correctly, then a risk rating finalised without the approval policy requires
- The complete twelve-scenario run, start to finish
Three sentences you can quote
Pre-cleared, attributable to Shaun Cuttill, CTO and co-founder:
“The agent didn’t get a new goal. It got better at pursuing the one it already had.”
“An action nobody has ever seen is still just an action asking for authority it may not have.”
“A control that the thing being controlled can turn off is not a control. It is a suggestion.”
What this is not
Stated plainly so you do not have to ask. The twelve-scenario run is our own demo environment, not a third-party audit and not a customer deployment. The separate ten-action comparison against NVIDIA OpenShell was run on lab hardware against a representative set of risky actions: it is not a benchmark, not a live incident and not a packaged integration. Mountain Theory is a member of the NVIDIA Inception Program, which is not a sponsorship, a partnership or an endorsement, and NVIDIA has not reviewed those results. Mountain Theory is patent pending. We publish what we do not stop as well as what we do.
Written before it happened
We wrote the failure mode down before it happened. The white paper below was published on 1 November 2024 by Michael May and Shaun Cuttill, and the Internet Archive holds an independent capture of it from 7 November 2025, which predates the coverage cycle of all three incidents above. We did not archive it ourselves and we cannot edit that copy.
- The 2024 white paper, on our site
- The same paper in the Internet Archive, captured 7 November 2025
- What we said then, against what happened since
An offer, if you would rather not take our word for it
Pick the scenario yourself. We will run it in both lanes, governed and ungoverned, and you publish whatever happens. Including if it goes badly for us. That offer only works because we already publish the runs we do not win.
Press and analyst contact: mike@mountaintheory.ai
Read the full drift run · Read the ten-action OpenShell comparison