Twenty-Five Dollars a Company

Status: Incident / Gambit Security, September 2026

Gambit Security, Palo Alto Networks Unit 42 and Anthropic (July to September 2026)


CISO Perspective: One operator ran three off-the-shelf agent tools against about 100 companies at about $25 a target and took more than 600,000 payment cards. When a model refused, the agents switched models. In the victim-side cases the attacker’s work ran through the victim’s own AI tools on stolen keys. A refusal at the model is not a control. The check has to sit on the action, whichever model is behind it.

  • Date. July to September 2026. Gambit Security published September 22; Semafor’s Unit 42 case September 16; Anthropic’s misuse report September 2026.
  • Type. Autonomous agent campaign against retail, hospitality and travel companies: payment card theft and collateral destruction. Plus victim-side intrusions run through the victim’s own AI tools.
  • What happened. 1,951 operator commands, 105 attack projects in five days, at least 27 companies compromised, more than 600,000 card records from two of them, 180 database tables dropped at one victim by a cleanup skill. Separately, a European software company’s own AI tools were turned against it in under 10 hours.
  • Where the gap is. Model refusals were routed around by switching models. On the victim side, a stolen key made the attacker’s agent indistinguishable from the company’s own, and nothing checked the action.
  • Why it matters to us. The action is the same whichever model issues it. The check belongs on the action, and it has to work with any AI the attacker or the victim happens to be using.

Gambit Security found the attacker’s staging server sitting open on the internet and read the whole campaign off it. One operator. Three off-the-shelf agent tools. Roughly 100 companies since July, at least 27 of them compromised in a single five-day stretch in September, and more than 600,000 payment card records taken from two of them. The model bill for the whole run was somewhere between $12,000 and $18,000. Per target, the mean was $25.46.

The operator typed 1,951 short commands in Chinese across 260 sessions, most of them an instruction to launch an attack. One tool found the weaknesses. One ran the exploitation, for hours at a stretch, on its own. One orchestrated the rest through the Hermes agent framework, loaded with 121 skills, 78 of them for attack. DeepSeek models did most of the work. When the newer models refused a request, the operator fell back to Claude Opus 4.6 through OpenRouter. Different vendor, different country of origin, same campaign.

The victims Gambit could name by type: a Fortune 500 hospitality company, a major US airline, an industrial supplies distributor, an online fashion retailer, a bicycle shop. Of the stolen cards, 79% belonged to US cardholders. A payment processor checked a sample and found at least 60% had never been flagged for fraud, which means the cards were fresh when they went to market.

Then there is the damage nobody ordered. One skill file told the agent to erase the card data at the source after taking it. At one victim the cleanup routine matched too broadly and dropped 180 database tables, including the backup tables the administrators had made for themselves. The agent was not told to destroy the company’s records. It was told to clean up, and it did what an agent does with a vague instruction and no check on the action.

Gambit put the conclusion plainly. “Remediation windows in complex environments are still measured in weeks,” and the attacker’s clock now runs in hours.

The Gambit agents belonged to the attacker. That matters, because the harder problem is the one where the AI doing the damage is yours.

Semafor published the anatomy of one such case on September 16, from Palo Alto Networks’ Unit 42. A European IT and software company, this summer. The attacker’s agents came in through a public API, pulled credentials that developers had left inside the software, took over the build and deployment system and the cloud keys, and then ran their operations “through the victim’s own AI tools.” Under 10 hours, against roughly two weeks for a human crew doing the same job. They left an extortion demand and an 80-page vulnerability report on the way out. Unit 42 got them out before what it called “full success.”

Anthropic’s misuse report for September says the same thing at scale. Stolen API keys, session tokens and devices have “increasingly become the sole objective of multiple criminal groups.” One breach “took only hours from first access to bulk data theft.” Another went from one stolen developer token to full cloud admin “in roughly three hours.” A Chinese espionage group ran agent swarms with a memory that persisted across sessions. A Russian group’s agents rebuilt their own malware when detection flagged it.

Together, those cases say the attacker no longer needs to bring an AI. Yours is already there, already authenticated, already allowed to touch the build system and the customer database. All they need is the key.

Andy Piazza at Unit 42 told Semafor that “even ill-intentioned AI agents need a human in the loop.” He is right, and the Gambit record shows what that human does: 1,951 times, someone typed a short instruction to launch. The loop on the attacker’s side is a person launching runs and reading results. Everything between those two moments happened at machine speed, on whichever model would take the job.

That is the part security teams should take from this. Most of what we built over the last 20 years assumes the thing doing the damage is either a person or a piece of code somebody wrote on purpose. An agent is neither. It reads its goal, picks its actions, and if one action is refused it tries another. The Gambit operator did not write an exploit for each of 100 companies. The agent worked each one out on the spot, then cleaned up after itself, and once it cleaned up too well.

There is a temptation in the Gambit story to make it about which models refused and which did not. Newer models refused. Opus 4.6 took the fallback work. DeepSeek did the bulk. The campaign ran on all of them.

A refusal at the model is a good thing and it is not a control. The attacker had a menu of models and switched when one said no. Inside your company the situation is the mirror image: your agents run on whatever model your vendor picked, and next quarter it will be a different one. A control that depends on the model’s manners is a control you renegotiate every time the model changes.

The action does not change. Dropping a table is dropping a table whether Opus wrote the command or DeepSeek did. Pushing a build with a stolen key is the same push. Reading 600,000 card records out of a database is the same read. If the check sits on the action, the model behind it is a detail.

Mountain Theory sits between the AI’s decision and its execution. Every action an agent is about to take is checked against rules written in plain English, before it runs. Allowed, held, or blocked. Every decision logged. It works with any AI you use.

I am not going to claim it would have stopped Gambit’s operator or the Unit 42 intrusion. I did not run them. What I can say is what the check is built to hold or stop. An agent that is compromised and starts taking the actions an attacker wants is, from where the check sits, an agent taking actions. A bulk read of payment records leaving your environment. A destructive command against a table the agent’s task never touched. A blocked action tried again by a different route. An agent that has drifted from the goal it was given to one nobody gave it. A run of retries past the cap. An action nobody has written a rule for yet, held for a decision instead of allowed through.

None of that asks which model is running. The model still decides. The action does not run until it has been checked. At $25 a company, the attacker’s cost of trying is close to zero. Make the cost of the action running without a check higher than that.

Sources. Gambit Security, “AI Agents Are Hacking Online Retailers for $25 a Company” (September 22, 2026). Forbes, Thomas Brewster (September 22, 2026). Semafor, J.D. Capelouto, “Anatomy of an AI-powered hack” (September 16, 2026), on the Unit 42 case. Palo Alto Networks Unit 42, “Chinese-Speaking Threat Actor Harnesses AI Models for Autonomous Cyberattacks” (July 30, 2026). Anthropic, Detecting and countering misuse of AI: September 2026. The stolen-access side of the same month is here.

The attacker picks the model. The victim does not. The action is the same either way, and the action is the only place the victim can put the check.

See it on your own agents. Book a demo.

Scroll to Top