Price per Million Tokens: The Complete Comparison of AI Coding Tools (Subscriptions, APIs, Aggregators, Local)

How much should a developer really spend on AI each month?


I measured my actual usage over fourteen days. I then applied the market's official pricing grids to calculate the price per million tokens for each offering. This time, I’m starting with the answer. The first version of this article gave all the figures, but did not clearly answer the question.

So here is the answer. The tables come afterward.

What you should get right away

Your situationWhat you getPer monthIn euros incl. VAT
You are learning or tinkering in the eveningGLM-4.7 Flash via API, free, and Ollama on your machine0 $0 € if you already have the graphics card
You code for a livingClaude Pro for judgment, GLM Coding Lite for volume, and a DeepSeek API buffer for overages58 $about 65 €
AI is your workstation, from morning to eveningClaude Max 20x, plus GLM Coding Pro as a second engine272 $about 303 €
There are five of youCopilot Business at 19 $ per seat, one shared GLM Coding Pro, and a pool of credits for testing190 $about 212 €, deductible if you have a company

And if you really don’t want to think about it, get Claude Pro for 20 $. You won’t get a bad deal. You’ll simply pay a little too much on the day you hit the limit.

There you have it. You can close the tab. The rest explains how I arrived at this answer. Above all, it shows why the price displayed by a provider has almost nothing to do with what you will actually pay.

The displayed price lies, and I can tell you by how much

A subscription no longer sells tokens directly. Claude Pro, ChatGPT Plus, and GLM Coding Plan sell five-hour windows and weekly limits. So to calculate a price per million tokens, you need to know your actual usage. I went and measured mine.

I analyzed the logs from my Claude Code sessions in ~/.claude/projects: 724 files, from July 28 to August 11, 2026. My script adds up the tokens reported by the API in each response, then applies Anthropic’s pricing grid.

What I measuredResult
Period14 days, from 07/28 to 08/11/2026, 724 session files
Tokens consumed5.85 billion
What the API would have charged me$4,867, or approximately $10,400 per month
Average cost per million$0.84, not $5 or $25
Cache reads95.81% of the volume
Cache writes3.60% of the volume
Output, the code actually written0.55% of the volume
Previously unseen input0.04% of the volume

Claude Opus 5 costs 5 dollars per million input tokens and 25 per million output tokens. Yet my actual cost comes down to 0.84 dollars. Why? A coding agent spends its time rereading the same repository. These rereads come from the cache, a copy of the context retained by the provider and billed at one-tenth the price. They account for almost all the volume. The code actually produced, meanwhile, accounts for only half a percent.

Comparisons that assume 70% input and 30% output arrive at 11 dollars per million with the Opus 5 pricing. I measure 0.84 dollars. Their result is thirteen times too high.

But another result made me sit up and take notice. I hadn't noticed it in the first version.

Deux barres empilees comparant la part du volume et la part de la facture pour Claude Opus 5

Three point six percent of the volume, twenty-seven percent of the bill. Cache writing is the popcorn of the movie theater.

Cache reads account for 95.81% of the volume, but only 57% of the bill. Cache writes weigh in at just 3.6% of the volume, but 27% of the price. A quarter of the bill for approximately one-thirtieth of the traffic. That's where your money goes. And nobody talks about it, because this cost does not appear clearly on any pricing page.

So I apply a simple rule throughout the rest: I recalculate every line using its four rates: input, cache reads, cache writes, and output. No universal formula. A model without caching bills every reread at the full price. The difference is brutal.

effectif = cache_lecture x 0,9581
         + cache_ecriture x 0,0360
         + sortie x 0,0055
         + entree x 0,0004

Two caveats. My usage is extreme: I orchestrate several sub-agents in parallel, which explains the twenty cumulative hours of session time per day in the measurement. Nor have I tested the forty offerings in this article. I use Claude Code every day and have tinkered with Ollama. For the rest, I rely on the official pricing.

The table, all at once

Here is the entire market in a single table, from least expensive to most expensive. The capability index is my own. I built it from the SWE-bench Verified and SWE-bench Pro results recorded on BenchLM on August 10, then supplemented it with my own assessment when no public score was available. An asterisk indicates estimates. Prices come directly from official sources.

ModelScoreInputCacheOutputEffectiveWhere to get it
GLM-4.7 Flash48*freefreefree0 $API and local
DeepSeek V4 Flash60*0,14 $0,0028 $0,28 $0,009 $API
DeepSeek V4 Pro74*0,435 $0,0036 $0,87 $0,024 $API, open weights
GPT-5.6 Luna58*0,20 $0,02 $1,20 $0,033 $API and subscription
MiniMax M264*0,30 $0,03 $1,20 $0,049 $API
Qwen3 Coder Next66*0,30 $0,03 $1,50 $0,051 $API, open weights
Qwen3.7 Plus68*0,40 $0,04 $1,60 $0,065 $API
GLM-4.770*0,60 $0,06 $2,20 $0,091 $API, open weights
Claude Haiku 4.562*1,00 $0,10 $5,00 $0,169 $API and subscription
GLM-5.276*1,40 $0,14 $4,40 $0,209 $API and subscription
Kimi K2.7 Code72*0,95 $0,19 $4,00 $0,239 $API, open weights
GPT-5.3 Codex82*1,75 $0,175 $14,00 $0,308 $API and subscription
GPT-5.6 Terra84*2,00 $0,20 $12,00 $0,330 $API and subscription
Gemini 3.1 Pro80*2,00 $0,20 $12,00 $0,330 $API and subscription
Claude Sonnet 5862,00 $0,20 $10,00 $0,337 $API and subscription
Grok 4.575*2,00 $0,30 $6,00 $0,393 $API and subscription
GPT-5.478*2,50 $0,25 $15,00 $0,413 $API and subscription
Kimi K378*3,00 $0,30 $15,00 $0,479 $API, open weights
GPT-5.6 Sol88*5,00 $0,50 $30,00 $0,826 $API and subscription
Claude Opus 51005,00 $0,50 $25,00 $0,844 $API and subscription
Claude Fable 59810,00 $1,00 $50,00 $1,687 $API and subscription
GPT-5.5 Pro85*30,00 $none180,00 $30,825 $API

Prices are shown in dollars per million tokens. They were recorded on August 12, 2026, from the official pages. The “Cache” column corresponds to the cache read price.

Graphique du cout effectif au million pour vingt-deux modeles, en echelle logarithmique

A ninetyfold price gap. A twofold capacity gap. Find the error.

Four observations immediately stand out. The last one corrects an error spotted by a very attentive reader: myself, two days later.

First, the cache read price determines almost everything. DeepSeek charges 0.0028 dollars for this read, or 2% of the input price. Most other providers charge 10%. For an agent that repeatedly rereads its context, this difference matters more than the displayed input price. DeepSeek V4 Flash thus falls below one cent per million, making it ninety-three times cheaper than Opus 5.

Next, two providers charge more than average for cache reads, and this almost goes unnoticed. Kimi K2.7 charges 0.20 times the input price for reads. Grok 4.5 charges 0.15 times, compared with 0.10 almost everywhere else. Their actual cost thus rises by 50% and 25%, respectively. 

Third, Gemini adds a trap that I have seen nowhere else. In addition to the read, Google charges for cache storage at 4.50 dollars per million per hour. If you leave the cache open during lunch, it continues to cost you money without rereading a single token. Pchhh. Now you know.

Finally, GPT-5.5 Pro. This model has no cache pricing at all. In the first version, I nevertheless applied a cache discount to it. A foolish mistake, with enormous consequences. Without caching, the 95.81% of rereads are charged at the full price: 30.83 dollars per million, or thirty-seven times Opus 5, not six times as I had written. It performs worse at coding and costs thirty-seven times more. For an agent, that is not expensive. It is disqualifying.

Two thresholds also deserve your attention. Gemini and Grok double their rates beyond 200,000 context tokens. For Grok, this doubling applies to the entire request, not just to the tokens exceeding the threshold. GPT-5.6 applies the same trap at 272,000 tokens. When the context crosses the limit, the bill does not rise gradually. It doubles all at once.

Subscriptions, and what they really hide

Here, we leave the public pricing tables and enter the fog. The table shows two things: the price of the offering and what the provider actually publishes about its limit. The second column already says a lot.

OfferPer monthWhat the limit really says
Claude Pro$20, or $17 annually5-hour windows, a weekly limit, no published figure
Claude Max 5x$100Five times Pro, still without a figure
Claude Max 20x$200Twenty times Pro. My measured result: 12.5 billion tokens per month, or $0.016 per million
GLM Coding Lite$1810,000 credits per week, published figure. Approximately 43 to 87 Mtok depending on your cache rate
GLM Coding Pro$7260,000 credits per week, published figure
GLM Coding Max$160140,000 credits per week, published figure
ChatGPT Plus$2015 to 90 messages per 5-hour window on Sol, with a credits-per-million grid now published
ChatGPT Pro$100 for 5x, $200 for 20xMultiples of Plus. Not unlimited, as the documentation makes clear
ChatGPT Business$25 per seat, $20 annually3,000 requests per week on the reasoning model
Google AI Pro€21.991,500 requests per day on Code Assist
Google AI Ultra€99.992,000 requests per day
SuperGrok$30A weekly quota expressed as a percentage. No absolute figure
GitHub Copilot Pro$10$10 in credits, consumed at the API rate
GitHub Copilot Pro+$39$39 in credits
GitHub Copilot Business$19 per seat$19 in credits per seat
Cursor Pro$20$20 in included usage, at the API rate
Cursor Ultra$200$400 in included usage
Zed Pro$10$5 in credits, then API rate plus 10%
Le Chat Pro$14.99150 fast responses per day

Three remarks:

GitHub Copilot switched to usage-based billing on June 1, 2026. You pay 10 dollars and get 10 dollars in credits. One for one. So it is no longer really a subscription, but an API surrounded by an editor. Same thing with Cursor and Zed, which resell you tokens at the editor's rate with a monthly allowance. The word “subscription” no longer means much for half of this table.

GLM remains the only one to publish its limit in verifiable units. Its system changed on July 30: prompts are no longer counted as “prompts” but as credits, with 10,000 credits per week for the Lite plan. My token equivalent assumes a 90.9% cache rate. That figure comes from the documentation, not from thin air. The other providers hide behind five-hour windows that are impossible to verify. Claude is the worst on this point, even though it is the one I use. I hate to write it, but it is true.

My own figure remains the most telling. Over fourteen days, I consumed the equivalent of 4,867 dollars in API usage. Compared with the price of Max 20x over the same period, that gives a ratio of 52 to 1. No subscription can be profitable with this profile. Either Anthropic is heavily subsidizing usage, or the limits prevent almost everyone from reaching this level. Probably both. I am not going to complain too loudly about it.

The four stacks, and when they crack

Quatre piles d outils cote a cote avec leur prix mensuel, de la gratuite a celle de l equipe

Four levels. The only choice is deciding which one you live on.

You are learning, $0. GLM-4.7 Flash via API does most of the heavy lifting. Ollama handles autocompletion and offline work. This stack cracks as soon as a task requires genuine reasoning: the model goes around in circles, and you need to know how to take back control. It only works if you already own a graphics card. Otherwise, budget 1,500 euros for a used 4090, or 42 euros per month amortized over three years. Free still comes with a bill.

You code for a living, $58. Claude Pro at $20 provides the judgment. GLM Coding Lite at $18 provides the volume. Around $20 of DeepSeek API usage absorbs the overflow. You have the cheap model lay the groundwork, then ask the best one to refine it. For 20 dollars, DeepSeek V4 Flash provides more than two billion tokens at the effective rate. Yes, two billion. This stack cracks if you spend all day in it. Claude Pro then reaches its limit before the others. That is the signal to move up to Max.

AI is your profession and you still program in the evening, $272. Claude Max 20x serves as the primary engine. GLM Coding Pro runs in parallel as a second engine. This is more or less my stack. The calculation is simple: my usage would cost 10,400 dollars per month through the API. Even dividing it by five to represent more reasonable usage, the subscription remains ten times cheaper. What do you lose? 272 dollars. At this level, you are mostly paying to stop watching the meter. Honestly, it is worth the price.

AI is your profession but you no longer program in the evening and on weekends: A Claude Code Max x20 subscription at $200, properly optimized (no huge Claude.md, Roslyn via MCP if you code in .Net (this avoids endless 'grep' commands)) could also be sufficient; it is up to you to assess your usage

There are five of you, $190. Copilot Business costs $19 per seat and covers autocompletion as well as repository integration. A shared GLM Coding Pro at $72 handles the major projects. Around twenty dollars in OpenRouter credits lets you test other models without opening five accounts. The drawback is that no one has permanent access to a cutting-edge model. If the whole team uses heavy agents all day, you need to move to Claude Team seats. The budget then changes by an order of magnitude.

Local models and aggregators, two detours worth knowing about

Running locally is not free. It is probably the most widespread lie on the subject. Electricity costs little: at 350 watts under load at Belgian rates, you get to around 0.55 euros per million tokens produced. The real cost is the graphics card. A used 4090 costs around 1,500 euros, or 42 euros per month over three years, before you have even written a line. A used 3090 at 700 euros runs the same models, around 20% more slowly, but cuts the amortization in half.

Local modelScoreVRAMThroughput on RTX 4090What I think
Qwen3-Coder 30B55*19 GB40 to 60 tok/sThe best choice. Its mixture-of-experts architecture activates only 3.3 billion parameters, making it fast despite its size. Its 256K context is enormous for local use. You need a 24 GB card
Qwen3.6 27B50*17 GBaround 30 tok/sA dense model, so all the parameters work on every token. It is nearly twice as slow for a comparable result
GLM-4.7 Flash48*19 GBnot measured on my systemA 198K context and the same parameter count. It is a credible solution, also available through a free API if you would rather keep your card for something else
Full Kimi and DeepSeek models70 and above*200 GB and abovevery slowGreat on paper, impractical on a developer's PC. You need a workstation equipped with several cards or a Mac with 128 GB, and the result is still sluggish. Best left to the curious

The capability gap is very real. A local model that fits into 19 GB of video memory scores around 55 on my benchmark, compared with 100 for Opus 5. This is not simply a difference in polish. It is the difference between an assistant that completes a line and an agent that refactors a module on its own. Local models are mainly useful for autocompletion, mechanical tasks, and code that must not leave any machine. For everything else, an inexpensive API costs less than the depreciation of the card.

Let's move on to aggregators. These are intermediaries that provide one account and a single bill for access to dozens of models.

OpenRouter does not add any markup to inference. Its documentation states this clearly. The service earns revenue from credit top-ups, with a 5.5% fee for credit card payments. Surprisingly, it can even be cheaper than the provider, because it routes the request to the least expensive provider in its network. Kimi K2.7 costs $0.68 there, compared with $0.95 at Moonshot. I had written, “for volume, go directly to the provider.” So you need to check on a case-by-case basis, because that is no longer always true.

Together AI also lists prices lower than those of some providers. Hugging Face announces zero markup and offers $2 in credits per month to PRO accounts. Fireworks publishes a pricing grid based on model size, from $0.10 to $1.20 per million. At Cerebras, the two Code plans, priced at $50 for 24 million tokens per day and $200 for 120 million, are listed as “sold out.” That's what happens when you sell speed.

What will change by autumn

Let’s start with some good news, which cancels out a sentence from my first version. I had written that Claude Sonnet 5’s introductory pricing would expire on August 31, before increasing to 3 and 15 dollars, and that I would update the table in September. The increase has been canceled. Anthropic has made the 2- and 10-dollar prices permanent. No update is therefore planned. Sonnet 5 remains at 0.337 dollars per effective million, giving it the best capacity-to-price ratio in the table.

DeepSeek, by contrast, announces a significant upcoming price increase in its documentation. Don’t build your annual budget around a rate of 0.009 dollars.

SuperGrok Heavy has also disappeared from xAI’s official pages. Only SuperGrok at 30 dollars and SuperGrok Plus at 100 remain. The 300-dollar amounts circulating online come from blogs, never from a page published by the provider.

My takeaways

A model’s listed price says almost nothing about your actual bill. The cache read price matters greatly, because reading the cache accounts for 96% of what an agent consumes. Cache writing matters too, since it generates a quarter of the bill for about one-thirtieth of the traffic. Two models listed at the same price can end up with a one-to-five ratio. As for a model without caching, it is eliminated immediately, regardless of its score.

The price gap between the top and bottom of the table is approximately ninetyfold, compared with only a twofold difference in capability. This does not make high-end models useless. If a small model fails at a task, its low price is of no benefit. Only the result matters. On the other hand, sending every task to the most expensive model is pure waste. A good stack therefore combines two models.

If you remember only one figure, remember this one: for 40 to 60 dollars per month, or approximately 45 to 65 euros including VAT, an independent developer can now access computing power that cost several thousand dollars a year ago. The question is no longer “is it worth it?” It is “which combination?” The answer is at the top of the page.

Article written on August 12, 2026; prices may change in the meantime.

Join the conversation

You need an account to comment on this article. Creating one is free and takes under a minute.

  • The XMLTV file, free to download every day
  • Comment on articles and reply to other readers
  • Get an e-mail when an article you follow is updated

No comments yet.

Une erreur s'est produite. Cette application peut ne plus répondre jusqu'à ce qu'elle soit rechargée.Veuillez contacter l'auteur. Reload 🗙