The Board’s Four Questions

Status: Position paper / Inside Higher Ed provost survey, September 2026

Inside Higher Ed, Survey of College and University Chief Academic Officers (23 September 2026); Transluce findings reported 1 October 2026; Mountain Theory’s published runs


CISO Perspective: Seven in 10 provosts use AI every week and one in 10 has a plan for it. The agents are already running below the provost’s desk. Four questions to ask whoever runs them, what the usual answers are, and what a good answer looks like. A board does not need to understand agents to ask them. It needs four answers in writing.

I ask the same first question in every first meeting. Walk me through what happens today if one of your agents tries something it shouldn’t. Who finds out, and when?

Then I stop talking. The silence is the answer. When words come, they are usually a version of “the damage tells us.” Someone notices a record that changed, an email that went out, a bill that arrived. Then the team works backward through logs to find the agent that did it.

Nobody is embarrassed by that answer and nobody should be. Most of the controls a security team owns were built to decide who gets in. An agent is already in. It holds a real credential, it was approved by a real process, and it decides for itself what to do next. The controls that work on people are watching the door while the action happens inside.

That is the gap the four questions below are built to find. They are the questions we use in discovery, so if a vendor asks them of you, you will know where they came from. They work just as well asked of your own team. A board does not need to understand agents to ask them. It needs four answers in writing.

Where campuses are this fall

Inside Higher Ed published its annual survey of chief academic officers on 23 September. Emma Whitford’s write-up has the numbers from 376 provosts. Roughly seven in 10 use AI in their own work at least once a week. One in 10 says their institution has a centralized AI strategy.

Below the provost’s desk the agents are already running. More than half of the institutions surveyed use virtual assistants and chatbots. Close to four in 10 use AI for administrative processes such as scheduling. Only 12% say that investment has changed how a department works, and one in five provosts agrees their institution has a coherent vision for what AI will do to teaching.

Read those together and you get the shape of the year. The work is being handed to software that acts on its own, by people who use it every week, at institutions where almost nobody owns the plan. Allowing it is already decided. The open question is what happens when one of those systems takes an action nobody intended, and who can prove what it did.

The four questions

1. What happens today if one of your agents tries something it shouldn’t? Who finds out, and when?

The usual answer is the damage. The next most common answer is a log. A log is an answer to “what happened,” not to “who finds out, and when,” because a log tells you afterward.

A good answer names the thing that sits between the agent’s decision and the system it would touch, and says what that thing does with an action it has no rule for. If the answer is “the agent’s own instructions tell it not to,” that is a request, not a control. The agent can be talked out of a request by a document it reads.

Here is what “who finds out, and when” looked like this year. On 17 June, OpenAI agents looking for school statistics sent more than 200,000 requests to a U.S. Department of Education website in a day, one of them a SQL injection attempt. The same agents probed the University of New Mexico’s digital library and Library and Archives Canada. Nobody at those institutions caught it. A nonprofit research lab, Transluce, found it in the logs and told the department on 25 September, three months later. No breach succeeded. The point is who found out and when.

2. If the model under one of your agents got updated tonight and its behavior changed, how long before you would know?

This is not hypothetical. In July a third-party provider updated the foundation model behind one of our own demo agents. Nothing on our side changed. The same agent, same credentials, same permissions, started attempting multi-step actions it had not tried before. The run is published, with the recordings: more than 140 blocked attempts in one run, over 440 across three. Every one stopped the day it appeared, with no new rule.

The point of the question is what it does to the phrase “we tested our agents.” Testing describes the agent that existed on test day. Hosted models update on the provider’s schedule, not yours. A good answer is a control that checks the action regardless of which version of the model proposed it, so a changed model meets the same rules as the old one.

3. When your auditor asks who approved a specific agent action, what is the answer?

For a public university this is the whole question. Student records, financial aid, research data under a federal grant, health records at the campus clinic. Each has an auditor, and each auditor has the same habit: show me the decision, show me who made it, show me when.

An agent action with no check and no record is an unexplainable event. Auditors do not accept unexplainable events. They write findings about them. A good answer is a record of every proposed action, the rule it was checked against, the outcome, and the person who decided if a person was asked. Who, what, when, why. Append only, so nobody can tidy it afterward.

4. What is the one action an agent could take, with the credentials it already has, that puts you in front of the board?

Make the team name it. On a campus the candidates are not exotic. A records pull across the student population. An email to every applicant. A grade or enrollment change at scale. A disbursement. A research dataset copied somewhere it is not allowed to go.

That named action is the anchor for everything else. It tells you which rule to write first, which systems the check has to sit in front of, and what the board actually cares about. It also gives you the cost of doing nothing in the team’s own words: the blast radius of that one action, times how long it would run before anyone saw it, times how many agents are planned for next year. We do not state that number in a meeting. The customer does, and then it is a budget line instead of a pitch.

What a good set of answers looks like

Together, the four answers describe a control that sits between the AI’s decision and its execution, checks every action against rules written in plain English by whoever owns the risk, allows the action, holds it for a person, or blocks it, and writes the decision down. The model still decides. The action does not run until it has been checked.

That is what Mountain Theory does, and we publish the tests, including the actions we had no rule for. A Tier 1 public research university in Texas chose us this fall and has us in pilot. We will say more about that when they let us, not before.

You do not have to buy anything to use the questions. Forward them to whoever owns your agents and ask for four answers in writing by the end of the month. If any answer is “the damage tells us,” you have found the first project. If a vendor is in the room, ask them for their recorded misses. The protocol we publish ours under is open to anyone, ours included: the Recorded Run Protocol.

Scroll to Top