
Somewhere in your procurement notes there is a Claude number you wrote down weeks ago. It is probably wrong, and not because anyone lied to you. The pages ranking for this query quote Sonnet 4, Haiku 3.5, Claude 3 Opus and Opus 4.6 at $15 input and $75 output. Six of the ten name a model Anthropic has already retired or superseded. The freshest one, published September 18, 2026, builds its entire break-even calculation on Sonnet 4 at $3 and $15 per million tokens, a generation that is no longer in the current line-up. So which rate is your budget actually built on? You already committed to an answer, whether you checked it or not. Which means the cost of a marketing team's Claude usage is either four dollars a month or forty, depending on a model name most guides get wrong. Both numbers are correct. Only one of them is yours.
Claude API pricing is metered per million tokens, and billed separately for input and output. As of September 30, 2026, the current line-up reads $1 input / $5 output per million tokens on Haiku 4.5, $2 / $10 on Sonnet 5, $4 / $20 on Opus 5.5 and $10 / $50 on Fable 5.1. Your bill is those two rates multiplied by two different token counts, which is why the headline number never matches the invoice.
Is Claude API pricing per token or per request?
Per token, on both sides of the exchange. Anthropic charges one rate for the tokens you send and a different, higher rate for the tokens the model writes back. There is no per-request fee and no monthly minimum, so a hundred small calls cost the same as one large call carrying the same volume.
That matters for marketing work, because a campaign brief and the copy it produces are not the same size and are not billed at the same price. A 2,000-word deliverable is roughly 2,700 output tokens. The brand voice document, the product facts and the instruction block you send with every single call are input tokens, and you pay for them again on every request unless you cache them.
Anthropic's billing rests on three units, and only three:
- Input tokens: everything you send, including system prompts, brand guidelines, retrieved documents and conversation history.
- Output tokens: everything the model generates.
- Cache tokens: input you have stored and re-served, billed at a discount rather than the full input rate.
At the current rates, output is exactly five times the input price across the entire line-up. Every model follows the ratio, from Haiku 4.5 at $1 and $5 up to Fable 5.1 at $10 and $50. That 5:1 ratio is the most useful single fact on this page, because it means your effective cost per token depends on how much the model writes back, not on the model tier alone.
One more definition before the arithmetic, because it changes the answer: the context window is the maximum the model can hold at once, up to 1 million tokens, and it varies by model. A long context is not billed as a long context. It is billed as input, every time you resend it. If you are comparing Claude against another provider, the same two-rate structure applies; our OpenAI API pricing breakdown runs the same arithmetic on the other side.
The current Claude rate card (verified September 30, 2026)

The table below was read off Anthropic's own model pricing documentation on September 30, 2026. These are USD rates per million tokens (MTok). If you are quoting this table anywhere, re-check it before you do: these rates move, and the entire reason this page exists is that other pages did not re-check.
Current models
Model | Input / MTok | Output / MTok | Cache read / MTok | 5-minute cache write / MTok | Batch input / output |
|---|---|---|---|---|---|
Fable 5.1 | $10 | $50 | $0.25 | $12.50 | $5 / $25 |
Opus 5.5 | $4 | $20 | $0.20 | $5.00 | $2 / $10 |
Sonnet 5 | $2 | $10 | $0.20 | $2.50 | $1 / $5 |
Haiku 4.5 | $1 | $5 | $0.10 | $1.25 | $0.50 / $2.50 |
Legacy models Anthropic still carries
You may be paying some of these rates right now without realizing it, because an integration pinned to an older model keeps working after the model is superseded. Opus 5, Opus 4.8, Opus 4.7, Opus 4.6 and Opus 4.5 all sit at $5 input and $25 output; Sonnet 4.6 and Sonnet 4.5 sit at $3 and $15. Opus 5 is no longer the current Opus. Opus 5.5 replaced it at a lower rate, which means a team that never touched its configuration is paying 25% more per token than the current tier asks for, on the same workload, for the same model family.
Three facts worth pinning down, since the query for this page returns so many guides that get them wrong:
- The current line-up is Fable 5.1, Opus 5.5, Sonnet 5 and Haiku 4.5. No page ranking for this query publishes above Haiku 4.5.
- Haiku 4.6, 4.7 and 4.8 do not exist. The most recent Haiku is 4.5. The version confusion runs in both directions: some guides write up a generation that was never released, and others count down to ones Anthropic retired in February 2026.
- The spread inside one family is ten times, on both sides. Haiku 4.5 costs $1 input and $5 output; Fable 5.1 costs $10 and $50 for the same million tokens. Nothing else on this page changes a bill as much as that choice does.
How Claude billing actually works (and which levers matter)
Four mechanisms change a real bill, and three of them are not on the pricing page at all.
Prompt caching is the biggest lever. You store input, then re-serve it at a fraction of the input rate. A cache read costs 0.1x the standard input price, so it pays for itself after a single read on the five-minute cache duration, or after two reads on the one-hour duration. Two exceptions are worth knowing: cache reads on Fable 5.1 run at 0.025x ($0.25 per MTok) and on Opus 5.5 at 0.05x ($0.20 per MTok). Everything else uses the standard 0.1x. A 5-minute cache write costs 1.25x the base input rate, a one-hour write costs 2x.
Batch processing takes a flat 50% off both sides. Nothing changes about the request except when you get the answer back. Confirmed per model in the current rates: Fable 5.1 at $5 and $25, Opus 5.5 at $2 and $10, Sonnet 5 at $1 and $5, Haiku 4.5 at $0.50 and $2.50.
Two modifiers stack on top of both. US-only inference (inference_geo) adds 1.1x across all token categories, cache reads and writes included. Fast mode, still a research preview on the Opus tiers, adds 2x. Anthropic's documentation states plainly that these multipliers stack with each other and with the batch discount. That stacking is the single most under-reported detail in this whole pricing discussion, and it is the difference between a cheap cached batch job and a mis-estimated budget.
Platform fees sit outside tokens. Managed agent sessions bill $0.08 per session-hour of active runtime. Web search bills $10 per 1,000 searches, with the retrieved content charged again as input tokens. Code execution includes 50 free hours per day per organization, then $0.05 per hour per container. For an agent-driven workflow, the runtime fee is a line item, not a rounding error.
What a marketing team actually pays: 2M input and 400K output tokens
Here is a workload you can recognise: content generation, SEO briefs and campaign copy for one month, totalling 2 million input tokens and 400,000 output tokens. No caching, no batch, nothing clever. The arithmetic is tokens multiplied by rate, divided by one million.
Model | Input cost (2M tokens) | Output cost (400K tokens) | Monthly total |
|---|---|---|---|
Haiku 4.5 ($1 / $5) | 2M × $1 ÷ 1M = $2.00 | 400K × $5 ÷ 1M = $2.00 | $4.00 |
Sonnet 5 ($2 / $10) | 2M × $2 ÷ 1M = $4.00 | 400K × $10 ÷ 1M = $4.00 | $8.00 |
Opus 5.5 ($4 / $20) | 2M × $4 ÷ 1M = $8.00 | 400K × $20 ÷ 1M = $8.00 | $16.00 |
Fable 5.1 ($10 / $50) | 2M × $10 ÷ 1M = $20.00 | 400K × $50 ÷ 1M = $20.00 | $40.00 |
Legacy Opus 5 ($5 / $25) | $10.00 | $10.00 | $20.00 |
Read that table once more, because it says something most people do not expect. At this workload, every model in the current line-up costs less per month than a single Claude Pro seat, and the cheapest costs less than a coffee. If you came here expecting API bills to be frightening, the marketing-copy workload is not where the fear is justified.
The workload is small for a structural reason: marketing work is high-input, low-output and repetitive. All three properties are exactly what caching and batching reward.
Take the caching case concretely, with the two numbers that decide it. Suppose your workload is 20 calls a month, each carrying the same 80,000-token brand and brief block as input, which comes to 1.6M tokens of stable input. Uncached on Sonnet 5, that input costs 1.6M × $2 ÷ 1M = $3.20. Cached, the first call writes at 1.25x ($0.20 for that block) and the 19 later calls read at 0.1x, about $0.016 each ($0.30 total), so the same 1.6M tokens cost $0.50. Your monthly total drops from $8.00 to roughly $5.30. That is what the payoff rule means in practice: the first read already pays for the write, and every call after the second is money you keep.
Now suppose you work alone, you write one new brand document every week, and each request carries a different context. Your first call writes at 1.25x and you may never read the block a second time, so you paid 25% more than plain input for nothing. Caching is a volume decision, not a feature to switch on reflexively. Batch is the simpler one: at a flat 50% off, it is free money for anything that can wait, and the two stack. A batched, cache-heavy Sonnet 5 month lands in the low single dollars.
If you want the deeper version of what the same workload costs on a Claude subscription, our Claude AI review walks the plans and the model access they include.
Subscription or API: where the break-even actually sits

A Claude Pro seat is $20 per month billed monthly, or $17 per month billed annually at $200 upfront. Max starts at $100 per month, with a 5x and a 20x usage tier. Team seats are $20 per seat per month billed annually, or $25 monthly, and a Team Premium seat is $100 annually, $125 monthly, which the vendor describes as five times the usage of a Standard seat. Enterprise is a different animal: $20 per seat per month billed annually, plus usage billed at API rates. That last one tells you something important. Even Anthropic does not treat a subscription and metered usage as opposites; the top tier is both, added together.
So the question is not which model wins. It is where your volume crosses the line. This is our calculation from the published rates, not a figure Anthropic publishes.
For transparency, our formula is: (input tokens ÷ 1M × input rate) + (output tokens ÷ 1M × output rate). A Pro seat at $20 per month buys:
Model | Pure input tokens bought for $20 | Pure output tokens bought for $20 | Blended tokens at a 5:1 input-to-output ratio |
|---|---|---|---|
Haiku 4.5 ($1 / $5) | 20,000,000 | 4,000,000 | 12,000,000 |
Sonnet 5 ($2 / $10) | 10,000,000 | 2,000,000 | 6,000,000 |
Opus 5.5 ($4 / $20) | 5,000,000 | 1,000,000 | 3,000,000 |
Fable 5.1 ($10 / $50) | 2,000,000 | 400,000 | 1,200,000 |
Every row uses the same arithmetic, so the blended column reduces to 12 divided by the model's input rate, in millions: at a 5:1 mix, six tokens of traffic cost ten input rates, which makes the blended per-million rate 1.67 times the input rate, and $20 divided by that lands on 12M, 6M, 3M and 1.2M respectively.
Here is the same comparison at three levels of actual use, on Sonnet 5:
Usage level | Monthly tokens (blended 5:1) | Metered API on Sonnet 5 | Pro seat ($20 monthly) | Cheaper option |
|---|---|---|---|---|
Light (a few briefs and posts a week) | 1M | $3.33 | $20 | API, by $16.67 |
Medium (a content operation running daily) | 6M | $20.00 | $20 | Break-even |
Heavy (agents writing, revising and researching in long sessions) | 45M | $150.00 | $20 | Pro, by $130 |
The formula column behind those rows: 1M blended costs $3.33, 6M costs $20, and 45M costs $150, all at a blended rate of $3.33 per million tokens on Sonnet 5.
The marketing workload above, at 2M input plus 400K output, is 2.4M tokens, which is under half of the break-even volume and therefore clearly an API case. The pattern is consistent: on current rates the API wins below roughly 6M blended tokens a month on Sonnet 5, the two options are equivalent around there, and above roughly 10M to 15M blended tokens the $20 seat wins on price alone, before you count any engineering time.
What the API-vs-subscription debate keeps missing is why heavy usage gets heavy. Agentic and coding sessions burn volume rather than value: a single long agentic session, the kind that reads a codebase or edits a dozen files, can consume 500,000 to 2 million tokens. A handful of those a day crosses the break-even inside a week, which is exactly the job the Max tier exists to do. Gartner's 2026 finding puts the mechanism at 5 to 30 times more tokens per task for agentic workflows than for standard chat. That is why a cheap rate card and a large bill are not a contradiction; they are the same sentence. If that is how your team works, read our notes on Claude Code for marketing before you build a metered budget around it.
Subscriptions win on three things, and none of them is marketing copy: agentic volume, predictability (a fixed line item with a five-hour rolling usage window rather than an uncapped meter), and no engineering cost (a seat needs no client, no key management, and no retry handling).
The cost-control checklist that actually changes the number
Ranked by measured impact, and each one tied to a mechanism rather than a slogan.
- Route by task, not by habit. The ten-fold in-family spread is the largest lever you own. Frontier models for final passes and judgment calls, a cheap model for drafts, summaries, metadata and classification. Independent research on intelligent model routing reports 40% to 85% cost reduction.
- Cache anything you send twice. Prompt caching cuts cost by 50% to 90% in published figures, and Anthropic's own multipliers agree. The condition is repetition: if the same block goes out on every request, cache it; if your context is different every time, do not.
- Batch everything that can wait. A flat 50% off both input and output, and it stacks with caching.
- Trim the context you resend. Conversation history is billed as input on every turn. A ten-turn exchange does not cost ten times one turn, it costs more than that, because turns one through nine ride along each time.
- Check your model names against the current rate card. This is the cheapest fix on the list and the most commonly missed. A configuration still pinned to a superseded tier pays more for the same work, and the tiers move without the integration noticing.
- Cap the runaway loop before it bills. Metered means uncapped. Put a spend limit in place before you automate anything that can call itself in a loop.
For context on why this matters at all: enterprise LLM API spend rose 36% in a single year, from an average of $63,000 per month to $85,500 per month, while per-token prices were falling. Unit prices fall, bills rise, because volume grows faster. That is the paradox this page exists to break. If you want the counter-position in another ecosystem, our DeepSeek vs ChatGPT comparison covers how a much cheaper headline rate changes what a workload costs.
Do you know which model your integration is actually calling? If you cannot answer that from memory, start there.
Frequently Asked Questions
How much does the Claude API cost per token?
Rates are per million tokens, verified September 30, 2026. Haiku 4.5: $1 input, $5 output. Sonnet 5: $2 / $10. Opus 5.5: $4 / $20. Fable 5.1: $10 / $50. The output rate is five times the input rate across the line-up, so a single figure never describes a real bill. Re-verify the table before committing a budget, because these rates move.
Is the Claude API cheaper than a Pro subscription?
For marketing-copy volume, yes, clearly. At 2M input plus 400K output tokens a month, metered Sonnet 5 costs $8.00 against a $20 Pro seat, and Haiku 4.5 costs $4.00. The break-even on Sonnet 5 sits around 6M blended tokens a month; below it the API wins, and above roughly 10M to 15M the seat wins. Agentic and coding sessions are where the seat takes over, because those consume 500K to 2M tokens per session.
What is prompt caching?
Storing input you send repeatedly and re-serving it at a discount instead of paying the full input rate. A cache read costs 0.1x the standard input rate, 0.025x on Fable 5.1 and 0.05x on Opus 5.5. A five-minute cache write costs 1.25x, a one-hour write 2x. The payoff rule: one read covers a five-minute write, two reads cover a one-hour write.
Which Claude model is cheapest to run?
Haiku 4.5, at $1 input and $5 output per million tokens, and by a factor of ten against Fable 5.1 at $10 and $50. That spread is on both sides of the exchange, so it compounds across a whole month. For most marketing tasks (drafting, summarizing, metadata, classification, first-pass copy) Haiku 4.5 is the right default, with a frontier model reserved for the final pass.
Do API credits expire?
Anthropic sells usage, not prepaid credit bundles, so there is nothing to expire. You are billed for tokens consumed, and unused capacity has no cash value because you never paid for it in advance. That is the practical difference between metered billing and a seat.
What is the difference between Claude Pro and the Claude API?
Pro is a fixed monthly seat with usage limits measured over a five-hour rolling window, no engineering work, and a client you do not have to build. The API is metered access billed per token with no cap, which you reach by writing or running something that calls it. Pro is predictable and capped; the API is flexible and uncapped. Team and Enterprise plans sit between them, with a seat fee plus usage limits or seat fee plus API billing.
Does the Claude API price include web search and agent runtime?
No. Those are separate platform fees: web search bills $10 per 1,000 searches with the retrieved content charged again as input tokens, and managed agent sessions bill $0.08 per session-hour of active runtime. Code execution includes 50 free hours per day per organization, then $0.05 per hour per container. Budget for them explicitly if your workflow uses them, because they do not appear in a per-token estimate.


