FOR AI ENGINEERS

YOUR CODE COMPILES. YOUR AGENT STILL ACTS.

You cannot unit test what a model will decide next week. Mountain Theory checks every action your agent proposes against rules written in plain English, before it runs, on any model and any framework. Agents run unmodified. Write the rule, ship the agent, and let the record show exactly what it tried.

The failure your tests cannot see

Static analysis tells you the code compiles. Unit tests verify functions on the inputs you gave them. Evals score model output on a fixed benchmark. None of that tells you what your agent does when it chains four tool calls together against live production data at 2am.

We found out the other way. This summer a vendor updated the hosted model behind one of our own demo agents, in a change aimed at making it more persistent at chaining tools. Nothing in our code changed. The agent started trying multi-step paths it had never tried before, with no release on our side.

What happened when the model changed underneath us

What we do

We control what autonomous AI does. The model still decides. The action does not run until it has been checked.

Every action gets one of three outcomes: ALLOW, HOLD, or BLOCK. HOLD is a rule you write for the actions that need a person, with the approver and the timeout declared in advance. Everything else runs at full speed. Every decision is logged with the rule that matched it.

The scenario

A vendor-hosted model driving our demo agent was updated overnight to chain tools more persistently, and the agent began attempting multi-step paths it had never tried. Every new attempt was stopped the day it appeared. No new rule, no signature, no patch, because the check evaluates what an action would do and whether the agent has authority to do it.

The July run, recorded, with the two actions we had no policy for

One rule, as you would write it

Policy

Reading application secrets is prohibited for every agent. No approval path. Log every attempt.

Plain English, no custom syntax, no code to change it. That rule is the one that stopped the secrets read in the NVIDIA OpenShell run before the credentials printed.

Rules live in one place and apply everywhere the agent runs. Change a rule and the change is live; nobody opens a pull request to do it.

Same 10 actions, three configurations, alongside NVIDIA OpenShell

It works with what you already built

Mountain Theory governs the action, not the API, so it does not care which model you call, which framework you built on, or where it runs. AWS, Azure, GCP, or your own hardware, including on-premises GPU infrastructure with no dependency on a commercial model provider. Agents run unmodified. You do not rewrite them or move them onto a different framework to be governed.

Delivered as an SDK inside your agent or as a service in your VPC. SDK access goes through the design partner program.

Every AI system we can govern, and the ones we cannot yet

What you hand your security team

An append-only record of every decision: the action proposed, the rule that matched, the outcome, the time, and who wrote the rule. It is produced by enforcement, not assembled afterwards, so the sign-off conversation gets short. Your security lead reads the record instead of your code.

What the record has to contain, and why the usual logs do not count

The ask

A 30-day paid proof of concept on your agents and your action set, not ours. Then one flat annual price for the site, with no per-agent count. Tests published, misses included, before anyone signs.

Book 30 minutes and bring the agent that is stuck in review

Every published test, on one page →

Proof

Four things with a published run behind them.

Tests published, misses included. The Recorded Run Protocol →

Scroll to Top