
Your agent's model bill has a ceiling by default. The server it deployed does not. Where do you turn on a hard cap?
Model APIs stop your agent at the tier ceiling. The cloud where it deployed a service usually just emails you. Where to set a hard cap, and where it fails.
In this article8
- A soft limit sends an email, a hard limit shuts the service off
- Where a hard spend limit exists, and where you only get a notification
- Model APIs: spend limits in OpenAI and Anthropic
- AWS: spend limits for new projects, budget actions for existing accounts
- Google Cloud Spend Caps: four services and a manual reset
- Vercel and Supabase: a switch you have to flip yourself
- What no ceiling will stop
- What to do before the agent gets a key
A hard spend limit shuts a service off when the bill reaches a set amount; a soft limit only sends an email. Today the hard ceiling that is on by default sits with model APIs: Anthropic ties it to your tier ($500 a month on Start), OpenAI has an approved limit per usage tier, and since July 23, 2026 any account can set its own hard limit on the organization and on each project. In the clouds where agents deploy what they build, you have to turn it on yourself, and every option comes with caveats. AWS launched a per-project spend limit on September 16, 2026, but only in a new experience rolling out to a limited set of customers. Google Cloud Spend Caps have been in preview since July 28 and cover four services. Vercel sends only emails by default; pausing is a separate switch. And every cap fires with a delay: OpenAI, Google and Vercel state plainly that you pay for the overage in between.
Nobody wants to wake up to an email sent at midnight warning about a budget limit and find that, while they slept, their rogue service had consumed several hundred (or several thousand) more dollars of usage.
Willison's post is short and makes one demand: every usage-billed service should ship with a hard spend limit on by default. You can remove it, but only by ticking an explicit box that says "I understand the service won't be shut off, and I pay for everything above the limit." His reason is direct: agents have made running code nearly free, and code that runs can spend money on paid APIs, hosting, storage and compute.
We checked how much of this already exists. The picture is odd. The default ceiling sits exactly where the agent thinks, in the model APIs. Where the agent ships what it wrote, there is no default ceiling anywhere, and turning one on comes with caveats the news summaries leave out.
A soft limit sends an email, a hard limit shuts the service off
Almost every cloud has had a "budget" for years. That is a soft limit: you hit the amount, you get an email. The service keeps running and the bill keeps growing.
Worse, the email is late. AWS Budgets data refreshes up to three times a day, typically 8 to 12 hours after the previous refresh. In that window a service an agent deployed and forgot gets half a day with nobody watching.
A hard limit works differently: you hit the amount, the service stops or the API starts returning errors. For a business that hurts, because the site is down. Willison's answer is that most companies and people would choose errors over a surprise $10,000 bill. We agree, especially for anything an agent spun up in an experimental project.
Where a hard spend limit exists, and where you only get a notification
| Hard ceiling by default | What stops | Main caveat | |
|---|---|---|---|
| Anthropic API | Yes, the tier ceiling | API returns 429 until the 1st | No separate limit on the Default workspace |
| OpenAI API | Approved limit per tier; your own via a toggle | API returns 429 | Not instant |
| AWS | No | The whole project is paused | New experience only, not available to everyone |
| Google Cloud | No | New calls to one service | Preview, 4 services |
| Vercel | No, emails only | Production deployments of all projects | Pausing is a separate setting |
| Supabase | Not checked; Spend Cap toggle on Pro | Usage above plan quota | Does not cover Compute |
Model APIs: spend limits in OpenAI and Anthropic
Anthropic sets the ceiling without asking you. The Start, Build and Scale tiers have monthly caps of $500, $1,000 and $200,000. Reach it and the API stops until 00:00 UTC on the first of the next month. You can set your own limit below the tier ceiling under Settings → Billing, in the Spend limits section. Through workspaces you can also give each group of keys its own limit.
OpenAI also has an approved monthly limit per usage tier, and since the week of July 23, 2026 a hard limit you set yourself is available to every account. You can set it for the whole organization and for individual projects.
- 1
OpenAI: organization
Organization limits → Spend → Edit spend limit. Enter a monthly amount and turn on Enforce a hard limit. Without that toggle the amount works as an ordinary notification.
- 2
OpenAI: project
Project settings → Limits → Spend → Edit spend limit, with the same Enforce a hard limit toggle. A project limit applies only to traffic billed to that project.
- 3
Anthropic: your own amount
Settings → Billing → Spend limits → Set limit or Adjust limit. You cannot go above your tier ceiling.
- 4
Anthropic: a separate workspace for the agent
Create a workspace, give it its own spend limit and issue the agent's key there. The Default workspace cannot have a separate limit.
The most useful part of both vendors' docs is the error codes. An agent that hits the limit gets a 429, and SDKs retry it by default. On Anthropic's tier ceiling there is no retry-after header, and retries won't help until next month. The only way to tell the ceiling from an ordinary rate limit is the code:
# Anthropic, tier ceiling: HTTP 429error.type = rate_limit_errorerror.details.error_code = enforced_spend_limit_reached# Anthropic, your own limit: HTTP 400error.type = invalid_request_error"You have reached your specified API usage limits..."# OpenAI: HTTP 429error.code = organization_spend_limit_exceedederror.code = project_spend_limit_exceeded
If your agent runs on a schedule, teach its handler to recognize these codes and stop instead of retrying. Otherwise you'll learn about the ceiling from logs holding a thousand identical errors.
AWS: spend limits for new projects, budget actions for existing accounts
On September 16 AWS announced a new builder experience, and it finally includes a limit that stops spending. It is set per project: AWS Settings → Billing → Cost by project → Set limit. When spending reaches the amount, AWS pauses the project and stops all its resources until the end of the month. Your data is kept.
There are more caveats than the announcement suggests.
- AWS offers the new experience through creating a new account, and the docs say outright that it is rolling out to a limited number of customers. Willison noticed this too and hopes for general availability.
- You need a paid plan. The minimum limit is the greater of $20 and a conservative estimate of your spending, which accounts for last month and running resources. You can't set $5 on a project that already has instances running.
- You can set limits on at most 10 projects.
- If a paused project sits untouched for 90 days, AWS deletes its data permanently.
On the plus side, there are early controls you enable separately, and they fire on forecasts. About 7 days before the limit, new resources stop launching; about 5 days before, idle resources are paused; about 4 days before, the top cost drivers among EC2, RDS, Lambda, Bedrock and SageMaker are paused. That last one is exactly for "the Lambda got stuck in a loop" and "Bedrock suddenly took off."
If you have a regular account and the new button isn't there, the closest working option is budget actions in AWS Budgets. You attach an action to a threshold: apply a deny IAM policy or SCP so nothing new gets created, or stop specific EC2 and RDS instances. This is not a ceiling. Lambda, storage and everything not listed keep spending, and the budget data itself, remember, lags 8 to 12 hours. Still, for a sandbox account where an agent experiments, it beats nothing.
Google Cloud Spend Caps: four services and a manual reset
Google announced Spend Caps on July 28, 2026. It is a new budget type: in Budgets & alerts choose Create new budget, then Spend cap enforcement instead of Alerts only, and pick the project, the service and the amount. You can't convert an existing budget into a spend cap, only create a new one.
The preview's limits are strict:
- One budget covers one project and one service, and the period is monthly only.
- There are four services: Gemini API, Gemini Enterprise Agent Platform (formerly Vertex AI), Cloud Run and Cloud Run functions.
- Only new calls are blocked. Requests already in flight finish and are billed. Persistent resources such as compute and storage keep accruing charges.
- Lifting the block is manual only, with a button in the console.
For AI services Google promises enforcement within minutes. For an agent that deployed a backend on Cloud Run and calls the Gemini API, that means two budgets, and it's worth creating both.
Vercel and Supabase: a switch you have to flip yourself
This is where the wording trap lives. In September 2025 Vercel announced that Spend Management is on by default for new Pro teams. That sounds like a hard limit. But the same changelog says deployments keep running without interruption unless you configure a hard limit manually. By default you get emails at 50%, 75% and 100%.
For the amount to actually stop spending, turn on Pause Production Deployments under Settings → Billing → Spend Management. When the amount is reached, Vercel stops production for every project on the team: sites, APIs and functions go offline until you re-enable each project. Checks run every few minutes, so Vercel itself advises setting the amount below what you're willing to lose. Spend Management is available on Pro and on Enterprise with Flexible Commitment; on the free Hobby plan projects already pause when the free allowances run out.
Supabase has a Spend Cap on Pro, managed in Cost Control on the organization billing page. It covers disk, egress, Edge Function invocations, monthly active users and storage. It does not cover Compute, read replicas, point-in-time recovery or IPv4: those are things you add deliberately, and you'll be charged for them regardless.
What no ceiling will stop
We read six sets of docs back to back, and three vendors say the same thing in different words: enforcement isn't instant, and you pay the overage. OpenAI writes that spending may slightly exceed the limit. Google advises setting the budget a bit below your absolute maximum and says outright that overage caused by the delay is on you. Vercel mentions a delay of a few minutes. The others are silent on it, and we wouldn't assume they behave differently. A hard ceiling limits the damage; it doesn't zero it out.
Where the summary is stronger than the docs
"AWS and Google Cloud launched hard limits" is true but incomplete. On AWS it's a new experience that isn't available to everyone yet. On Google it's a preview for four services where persistent resources keep billing. Before you count on a ceiling, open your own console and check that the button you need is actually there.
Second, a ceiling protects you from the bill, not from what the agent did. It won't recall an email blast to other people's addresses or close access the agent opened to the outside. We showed what it looks like when an agent gets out of its sandbox in our breakdown of the DNS incident in OpenAI's sandbox.
Third, specific to AWS: a paused project has to be unfrozen within 90 days. The limit that saved you from a bill can cost you your data three months later if you forget about it.
What to do before the agent gets a key
Our opinion, open to argument: put the ceiling not on the whole account but on a separate place where the agent lives. We're wrong if you have a single production project and the agent works only there: a separate place gives you nothing, and the real question is whether you accept the site going down at the ceiling.
- A separate project or workspace for the agent. A project in OpenAI, a workspace in Anthropic, a project in AWS or Google Cloud. The agent's key is issued only there, and that place's limit is lower than what you're willing to lose overnight.
- A ceiling below your real maximum. Limits fire with a delay. If you can afford to lose $100, set $70–80. Size it from the cost of a task, not of a token: why those diverge is clear from our comparison of GPT-6 Sol and Opus 5.5 by cost per task.
- Code that stops at the ceiling instead of retrying. The error codes are above; handling them takes a few lines.
- An agent that picks providers with a ceiling. This is Willison's own idea: agents could recommend services with hard limits and warn newcomers about services without them. Until that happens, you can write the rule into your project instructions, for example in CLAUDE.md: "don't deploy anything paid without a spend limit, and ask first."
A budget at the orchestrator level, where the agent gets a sum per task, is a different layer of protection with its own gaps: we showed where it reacts too late in our Paperclip breakdown. The provider's ceiling fires last, which is why you always need it.
Versions and prices as of October 4, 2026
Anthropic tier ceilings, console paths and the list of Google Cloud Spend Caps services are taken from vendor documentation as of October 4, 2026. AWS spend limits and Google Spend Caps are still rolling out or in preview: the interface and limits will change. Check your own console before relying on a limit.
Sources12expand
- Simon Willison, "We're going to need default hard budget caps on pretty much everything", October 3, 2026 — https://simonwillison.net/2026/Oct/3/default-hard-budget-caps/
- AWS, "New AWS experience helps builders get started and ship faster", September 16, 2026 — https://aws.amazon.com/about-aws/whats-new/2026/09/New-AWS-Builder-Experience/
- AWS Docs, "Create a spend limit in AWS Settings", checked October 4, 2026 — https://docs.aws.amazon.com/accounts/latest/reference/create-spend-limit.html
- AWS Docs, "Managing your costs with AWS Budgets" and "Configuring budget actions", checked October 4, 2026 — https://docs.aws.amazon.com/cost-management/latest/userguide/budgets-managing-costs.html
- Google Cloud, "New early anomalies and spend caps on Google Cloud Budgets", July 28, 2026 — https://cloud.google.com/blog/topics/cost-management/new-early-anomalies-and-spend-caps-on-google-cloud-budgets
- Google Cloud Docs, "Spend cap budgets", checked October 4, 2026 — https://docs.cloud.google.com/billing/docs/how-to/budgets-spend-caps
- OpenAI, "Spend limits", checked October 4, 2026 — https://developers.openai.com/api/docs/guides/spend-limits
- OpenAI Developer Community, "Hard spend limits rolling out to all API Platform accounts", July 23, 2026 — https://community.openai.com/t/hard-spend-limits-rolling-out-to-all-api-platform-accounts/1387914
- Anthropic, Claude Docs, "Rate limits", Spend limits section, checked October 4, 2026 — https://docs.claude.com/en/api/rate-limits
- Vercel Docs, "Spend Management", updated September 18, 2026 — https://vercel.com/docs/spend-management
- Vercel, "Spend Management now enabled by default on Pro", September 9, 2025 — https://vercel.com/changelog/spend-management-now-enabled-by-default-on-pro
- Supabase Docs, "Control your costs", checked October 4, 2026 — https://supabase.com/docs/guides/platform/cost-control
Read next
Agents now get a boss, a task queue and a budget. Do you need one if you already work in Claude Code?September 26, 2026
The alarm went off after 12 minutes. The agent was stopped two and a half hours later. What actually broke in OpenAI's sandbox?September 27, 2026
GPT-6 Sol is half the price of Opus 5.5 per token. So why is a task only 20% cheaper?September 26, 2026
Comments