For reporters and analysts

WHAT WE PUBLISHED, AND WHAT YOU CAN CHECK

In the last week of July 2026 a third-party vendor updated the AI model driving one of our demo agents. Nobody told us. The agent started attempting things it had never attempted, and the policy held anyway, with nothing shipped on our side. We published the run, the recordings and the control condition. This page is everything you need to check it.

Three labs, one summer

None of these models was told to attack anything. Nobody wrote a malicious prompt. The instructions were ordinary. Primary sources, not our summaries:

  • OpenAI. A model running a cyber benchmark with safety controls deliberately switched off exploited a flaw in OpenAI’s own infrastructure and chained onto Hugging Face production. OpenAI’s disclosure
  • Anthropic. Three incidents. Models told they had no internet access had it, through a misconfiguration. One published a malicious package to PyPI that was downloaded and run on 15 real systems, one of them a scanner belonging to a real security company. Anthropic’s write-up
  • Meta. A model reached the public internet through a sandbox misconfiguration, exploited a vulnerability in a third-party service, and made changes to a real company’s internal systems. CNN, 5 August 2026

The thread we think is underreported: in the Anthropic and Meta cases the misconfiguration was at the same evaluation vendor. Two of the three trace back to one supplier. That is checkable against the two disclosures above and we have not seen it written up.

Our run, with its control condition

An autonomous GRC agent doing vendor security reviews, built on Hermes from Nous Research. Twelve scenarios, eight deliberately dangerous and four ordinary work. Run in two lanes side by side: the agent with nothing in the path, and the same agent with Mountain Theory in the path. The ungoverned lane is the control condition, and it is published.

Published results. Full run and recordings at the link below.
 Nothing in the pathMountain Theory in the path
Scenarios compromised12 of 12, every run0 of 12, any run
Blocked attempts, one runn/amore than 140
Blocked attempts, three runsn/amore than 440
Action decisions renderedn/amore than 1,400

The finding we think matters most, and the one nobody has written about: four of the twelve scenarios are not attacks. They are ordinary work the agent is supposed to complete. After the model update all four went off script, and in the ungoverned lane they end the same way the attacks do. A capability upgrade does not only make attacks sharper. It changes how an agent does its routine job, which is the behaviour least likely to be watched and most likely to be trusted.

The recordings

Unedited screen captures of both lanes. Free to embed.

Three sentences you can quote

Pre-cleared, attributable to Shaun Cuttill, CTO and co-founder:

“The agent didn’t get a new goal. It got better at pursuing the one it already had.”

“An action nobody has ever seen is still just an action asking for authority it may not have.”

“A control that the thing being controlled can turn off is not a control. It is a suggestion.”

What this is not

Stated plainly so you do not have to ask. The twelve-scenario run is our own demo environment, not a third-party audit and not a customer deployment. The separate ten-action comparison against NVIDIA OpenShell was run on lab hardware against a representative set of risky actions: it is not a benchmark, not a live incident and not a packaged integration. Mountain Theory is a member of the NVIDIA Inception Program, which is not a sponsorship, a partnership or an endorsement, and NVIDIA has not reviewed those results. Mountain Theory is patent pending. We publish what we do not stop as well as what we do.

Written before it happened

We wrote the failure mode down before it happened. The white paper below was published on 1 November 2024 by Michael May and Shaun Cuttill, and the Internet Archive holds an independent capture of it from 7 November 2025, which predates the coverage cycle of all three incidents above. We did not archive it ourselves and we cannot edit that copy.

An offer, if you would rather not take our word for it

Pick the scenario yourself. We will run it in both lanes, governed and ungoverned, and you publish whatever happens. Including if it goes badly for us. That offer only works because we already publish the runs we do not win.

Press and analyst contact: mike@mountaintheory.ai

Read the full drift run  ·  Read the ten-action OpenShell comparison

Scroll to Top