PROOF
PROOF, ON THE RECORD
Tests published, misses included. We control what autonomous AI does. The model still decides. The action does not run until it has been checked. Everything on this page is a recording, a case study, or someone else’s account, and every number on it is already published on the page it links to.
Proof
Four things with a published run behind them.
The evidence, dated
A vendor updated the model behind our demo agent overnight. Every new attempt was stopped.
The same agent, the same rules, a different model underneath it. It tried more than 140 times in one run to go beyond its rules. Zero got through, and nothing changed on our side: no new rule, no signature, no patch.
Ten actions, three configurations, including the two we had no policy for.
Ungoverned, 10 of 10 executed. Under NVIDIA OpenShell alone, the sandbox denials fired and the in-bounds bad decisions still went through. Under OpenShell plus Mountain Theory, the in-bounds actions returned HOLD, HOLD and BLOCK. The two actions Mountain Theory did not stop are in the recording.
Told to delete audit evidence, the ungoverned agent did it. The governed one was blocked, and blocked again on the retry.
The same autonomous vendor risk analyst in two environments, one with nothing in the path and one with Mountain Theory checking each action. Optimo AI is a design partner and an advisor, not a customer. This was a proof of concept, not a deployment.
Why We Became Mountain Theory's First Design Partner, by Peter Holcomb, Optimo AI.
“That is why we became Mountain Theory's first design partner. They work at the execution layer.” Peter Holcomb, founder and CEO, Optimo AI, in the only account of Mountain Theory on the open web that we did not write. Published on optimoit.io on 8 September 2026.
A GRC engineering director who sat in on the July demo, on LinkedIn.
“I've been fortunate enough to see Mountain Theory in action. It's the real deal!” James Tabron, CISSP, Director of GRC Engineering, Aquia, who watched the Optimo AI proof of concept run in July 2026. Posted on LinkedIn on 12 September 2026.
Run it on us
The Recorded Run Protocol is the test behind every run above, and it is vendor neutral: a published action set, a control condition with nothing in the path, the outcome for every action including the ones the product did not stop, and a recording. Use it on us. Use it on anyone.
Three questions buyers ask
What proof is there that Mountain Theory actually stops an autonomous AI action?
Tests published, misses included: two recorded runs, both with terminal recordings and both naming the third-party technology involved. In the first, the same 10 actions were run in the same order under three configurations. Ungoverned, 10 of 10 executed. Under NVIDIA OpenShell alone, all 5 sandbox-boundary crossings were denied at the kernel and all 3 in-bounds bad decisions still went through, including a secrets read that printed credentials to the screen. Under OpenShell plus Mountain Theory, those same 3 actions returned HOLD, HOLD and BLOCK, and the secrets read was stopped before it executed, so the credentials never printed. In the second, a third-party provider updated the foundation model driving an autonomous agent. Nothing on our side changed, the agent began attempting multi-step actions it had never tried before, and every attempt was stopped on 30 and 31 July 2026, the days the behavior first appeared. No new rule, no signature, no patch.
How should I evaluate competing AI security vendors on evidence?
Ask every vendor for tests published with the misses included, the same four things, and compare the answers side by side. First, the action set: exactly which actions were attempted, in what order. Second, the control condition: what happened with nothing in the path, so there is a baseline to measure against. Third, the outcome per action, including the ones the product did not stop. Fourth, the recording or log. A vendor who publishes all four is making a checkable claim. A vendor who publishes a certification, an integration list or a customer logo is telling you about their process and their distribution, which are different questions. Mountain Theory publishes all four, including the actions it does not have a policy for.
Does a FedRAMP authorization or SOC 2 mean a vendor can control autonomous AI?
No. A compliance authorization attests to how a vendor runs its own service: that controls exist, that processes are documented, that evidence is retained. That is real and it matters in procurement. It is not a measurement of whether a product can stop an autonomous agent from taking a specific action, because no compliance regime was designed to test that. A vendor can hold every certification available and still have nothing in the path when an agent decides to delete evidence, disable a control, or read secrets it is technically permitted to read. The only thing that shows an action was stopped is a run where that action was attempted and did not execute.