On September 29, one of our own AI coding agents made between 14,000 and 15,000 calls to Anthropic’s API in at most 95 minutes. Shortly before midnight Berlin time, Anthropic disabled our organization, and every feature we run on Claude stopped at once. Nothing was hacked, and nobody meant any harm. That’s the uncomfortable part.
This is the write-up: what the agent did, why it could, what it actually cost, and what we changed. It’s a little embarrassing to publish. That’s also why we think the advice at the end is worth your time.
The harmless sentence
That same morning we had moved all our production AI usage into a new Anthropic organization, with one workspace per site and one API key per job. We funded it with $100 in prepaid credits. It was a tidy setup, and I was pleased with it.
That evening, a Claude Code session was working on a phone assistant for a photo studio that answers callers’ questions from the studio’s own knowledge base. The session could hand work to sub-agents, and its permission prompts were off. Then I gave it this instruction:
“I want every question to generate the one best answer, I don’t care how it’s done.”
Read that the way an agent reads it. There’s no budget in it, no deadline and no limit on method. I meant “don’t bother me with the details.” The agent heard a scope.
What the agent did with it
It did good engineering, which is what makes this worth telling. It built a topic router: Claude reads the caller’s question together with the full list of topics in the knowledge base and picks the one that fits. To find out whether the router was any good, it needed test data. Its sub-agents wrote 2,750 realistic caller questions about a knowledge base of about 100 topics, including 282 it can’t answer, where the right response is to say so. Two independent AI judges labeled the right answer for each; where they disagreed, a third judged blind. That part ran on the Claude subscription we develop with, at no API cost.
Then it measured. It wrote an evaluation script that sent the questions through the router via the API and compared the answers with the labels, then adjusted the routing rules and ran everything again. Every tweak re-ran thousands of questions across Haiku, Sonnet and Opus, with up to 16 calls in flight at once.
- Between 14,000 and 15,000 API calls in at most 95 minutes, between 10 p.m. and the shutdown.
- About 50,000 tokens per call, because every call carried the entire topic catalog.
- About 720 million tokens in total, 717 million of them read from the prompt cache.
- $152.64, almost all of it on Sonnet and Opus.
- The result: on two question sets the router had never seen, Opus alone picked the right topic 91.1% of the time on the first, and the setup we’d run in production, Opus with Sonnet as a timed backup, 90.9% on the second. The keyword matcher it would replace scored 42.9% and 49.7%.
The router works. At 11:35 p.m. Berlin time, the API began refusing every request from every one of our keys, at first with this line:
This organization has been disabled. An organization admin can appeal at https://console.anthropic.com/appeal
The agent did the right thing at that point. It stopped and wrote itself a note: “never run bulk evals on the production key.” Correct, and about an hour and a half late.
Why it could
Three things had to be true at the same time. All three were.
- A live key sat in the project folder. The phone assistant is server code, and server code reaches Claude through an API key. For local testing, that key lived in the project’s
.envfile, as it does in countless projects. - The code loaded it by itself. The agent never went looking for a key. It ran its evaluation script, the script read
.envthe way the project’s code always does, and every request was billed to the production organization. - Nothing asked. With prompts off, there was no moment where a person saw “about to send thousands of paid requests” and could say no. I knew tests were running; late that evening I asked the session to check on them and keep going. I didn’t know how big they were or what they cost. Auto-reload didn’t ask either: at 11:02 p.m., with the first $100 almost gone, it bought another $84.58 in credits by itself.
The second point is the one to remember. Claude Code can block its own file tools from reading .env, and that rule is worth having. But Anthropic’s documentation says plainly that such deny rules don’t reach a script that opens files on its own—in its words, “a Python or Node script that opens files itself.” That is exactly what happened here. A deny rule on the agent can’t stop the code the agent runs. Anthropic’s documentation names what can: Claude Code’s sandbox, which blocks a file for every process a command starts, once you list the file there and turn off the option that lets a blocked command retry outside the sandbox. A key that isn’t there needs no setting at all.
Why the API, when the subscription had paid for everything else? Because the thing under test was server code, and in production, server code calls Claude with an API key, not through a developer’s subscription. Measuring the router honestly meant calling it the way production would. The agent’s reasoning was sound. The key the code picked up was the wrong one.
What it really cost
The $152.64 is annoying, not dangerous. The prompt cache kept the cost low; it didn’t make the traffic any smaller. The real bill arrived at 11:35 p.m. When an organization is disabled, every workspace and every key in it stops—including the ones that had nothing to do with the test.
For us, that meant every feature we run on Claude went dark at once, on our own site and on the client sites that use them:
- answers in site search,
- automatic replies to contact form inquiries,
- translation of new content,
- and the scans and audits behind Cited, our AI visibility tracker.
That’s the lesson we didn’t see coming. We had split our sites into workspaces so that one site’s usage could never eat into another’s budget. With a spend limit on each, workspaces do that well; ours didn’t have limits yet. And even with limits, they do nothing against a decision about the whole organization. The structure we built to contain risk had put everything behind one switch.
Anthropic didn’t tell us why at the time, and we won’t guess in public. We can describe what the account looked like that evening: an organization less than a day old, a sudden burst of parallel traffic, hundreds of millions of tokens in an hour and a half. Anthropic’s documentation says new organizations may start on lower limits while they build a usage history, and it recommends ramping traffic up gradually. We did the opposite on day one. We filed an appeal on September 30.
The guardrails we run now
None of this needs new technology. It needs the defaults we should have set on day one, ordered from the ones that work without anyone’s cooperation to the one that doesn’t.
No API keys in our development .env files
No .env file in the project folders we develop in holds a key or token in plain text anymore. They moved to our password manager, a scan of every one of those files found none left behind, and both development keys are deactivated. An agent can’t reach for a key that isn’t there, and neither can the code it runs.
Spend limits per workspace, a ceiling for the organization, auto-reload off
Every workspace that holds a key now has its own monthly spend limit, and the organization has one above them. The Claude Console offers both; a workspace limit can only be set lower than the organization’s. One detail worth knowing: the Default Workspace can’t carry a limit, so ours holds no keys at all.
A spend limit also stops things when it’s reached. The difference is scale and control: a workspace limit stops one workspace, and we can lift any of our limits ourselves in a minute. An appeal takes as long as it takes. Auto-reload goes off too. On September 29 it was on, and it topped the balance up in the middle of the run. A limit that refills itself isn’t a limit.
A spend alarm that checks twice an hour
A check on our server reads Anthropic’s usage and cost reports twice an hour. It warns us when a workspace’s spend climbs fast, when a day’s spend passes a set amount, or when a workspace nears its limit, and it raises the alarm when a workspace hits its limit or the API stops accepting our key. It warns; it doesn’t stop anything. Only a spend limit does that.
No paid API calls without an OK
Our agents work under a standing rule, written into the instructions file that every agent session in our project folders loads: more than 100 calls to any paid API, or an expected cost above $5, needs my explicit OK, with the number of calls and the cost stated up front. The same file says production keys belong on the server, not on development machines. It’s the weakest guardrail on this list, because it depends on the agent following it. That’s why it comes after the three that don’t.
Evaluations that are sized before they run
- Do the token math first. Calls × tokens per call × repetitions, then the price. For this router, one pass over 2,750 questions on three models at 50,000 tokens each is about 410 million tokens—before the first rule tweak. Anthropic’s token-counting endpoint is free, so there’s no reason to guess.
- Test on 200 questions, not 2,000. A small sample shows whether a change helps. The full set runs once, at the end.
- Don’t send the whole catalog with every call. Narrow the candidates cheaply first, with a keyword or embedding search, and let the model choose among a handful of topics instead of about 100.
Six things to check today
If you use AI coding agents, or your website calls an AI API, these take about an hour together.
- Find the keys. Search every folder an AI agent works in for
.envfiles and config files that hold API keys. Any live key there is a key the agent can use. Move it out. - Set spend limits on the organization and on every workspace. In the Claude Console, the organization limit is under Settings > Billing, and each workspace has its own Spend limits tab. Keep production out of the Default Workspace, which can’t have one. If you use other AI providers, look for the equivalent settings before you need them.
- Switch auto-reload off. It’s on the same Billing page. An empty balance stops a runaway job; a balance that refills itself doesn’t, and then only a spend limit stands in the way.
- Check your agent’s permission mode. If prompts are off, run the agent somewhere without keys. Anthropic’s own guidance for Claude Code’s bypass mode is an isolated container or VM. Add deny rules for
.envtoo, and switch on Claude Code’s sandbox with the file blocked and unsandboxed retries turned off, so the block also holds against a script that opens the file itself. - Size bulk runs before they start. Calls × tokens × repetitions, then the price. Count tokens for free, test on a small sample, and send work that can wait through the Batch API, which Anthropic bills at half price.
- List what stops when the key stops. Write down every feature that depends on a single AI provider and decide what each one does when the API refuses: hide itself, fall back or say so. We learned our list at 11:35 p.m.
How it ends
On October 1 at 10:23 a.m., about 35 hours after the shutdown, Anthropic’s Safeguards team wrote back. Their mail said the account had been disabled for violating the Usage Policy; after our appeal, they had “investigated this decision and reinstated your account,” and they apologized. They didn’t say which part of that night’s traffic tripped it, and we still don’t know.
The router itself was parked for good. Since October 1, the phone assistant picks its answers with a model running on our own server: right 83% and 79% of the time on two fresh test sets, with no paid API behind it.
Two things we’d tell anyone in the same position. File the appeal right away; the link is in the error message itself. And don’t wait for the answer to start fixing the cause: we began the next morning. The guardrails above stay. They should have been there before the first key was created.
If you run AI features on your own site and want someone to go through keys, limits and fallbacks with you, send us a note.