[BidClub_]
SemiAnalysis · · 50 min

Ep. 036 - $200 Buys $12,000 of Opus Tokens, We Bought Every Plan (Tokenomics)

Max KanJordan NanosAndrew Megalaa

AI & SoftwareTechnicalCompany Building
YouTube ↗
TL;DR
  • Anthropic’s $200 subscription can deliver roughly $12,000 of usage at API prices without necessarily losing money. Max cites a consensus forecast of more than 90% API gross margins; if the average subscriber consumes only about 10% of the maximum allowance, subscriptions could still earn 50–60% gross margins. The sharper conclusion is that API pricing may be “punitive, exploitative, speculative pricing,” not that every subscription token is subsidized.

  • A subscription’s value cannot be reduced to one dollar figure because each model and token type consumes plan credits differently. Fresh input, cache writes, cache reads, and output carry different weights, so the same $200 plan can produce radically different API-equivalent values depending on the workload. The defensible formulation is always: “a certain workload on a certain model” yields a particular value.

  • Caching is the decisive economic lever in agentic inference. The tested agent workload was roughly 96% cache reads; those reads can cost 0.1x or less than fresh input because they retrieve an existing KV cache instead of repeating model computation. Maintaining high cache-hit rates across deployments, memory tiers, and pauses is difficult systems work—and “exactly what distinguishes good inference providers from worthless ones.”

  • Anthropic offered materially more usable tokens than OpenAI in the speakers’ like-for-like $200 comparison. They measured about 1.9 billion Astra tokens versus 3.1 billion Fable 5.1 tokens, with Fable consuming only 50% of the Anthropic allowance and leaving the other half for Opus. Although critics cited an Artificial Analysis comparison of roughly $5 for a maximum-cost Opus 5.5 task versus $0.72 for a GPT-5.1-Codex task, Andrew Megalaa argues that matching benchmark settings narrow the efficiency gap to 2–3x while Anthropic supplies about 5x the API-equivalent value.

  • The industry advantage accrues to model owners, while wrappers inherit structurally worse economics. Anthropic can monetize API traffic at greater than 90% gross margin, whereas services such as Perplexity or Cognition may have only around 20% of the API cost available to cover their own economics and consequently offer much lower limits on Anthropic models. The episode’s investor framing is blunt: in AI, possessing a model “that everyone wants to use” is a huge advantage, while being a “wrapper” is especially unattractive.

  • Realized subscription margins depend on ordinary users consuming nowhere near the published maximum. Max believes Anthropic utilization “cannot be much higher than 10%” if leaked financial figures are to reconcile, implying approximately 40–70% subscription gross margins. OpenAI’s announced 50% limit reduction had not yet hit users who bought plans before the announcement because of a one-month grace period, potentially postponing backlash until the limits become tangible.

  • Open models can win large volumes of adequate work without invalidating frontier-model economics. Businesses can rationally assign simpler, margin-sensitive tasks to open models such as GLM-5.3; Jordan Nanos did so when frontier-model cybersecurity refusals blocked his workflow, then reused the same context for a trivial ten-line PR at roughly one-tenth the tokens. The unresolved frontier thesis is whether the economy can profitably absorb the equivalent of “100 million super-smart PhDs” faster than routine work migrates to cheaper models.

Digest · the substance, structured for research

1. The $12,000 headline reflects API pricing, not proof of an $11,800 subscription loss

  • Max’s framing: paying $200 and receiving $12,000 of API-priced Claude tokens is unquestionably “a good deal,” but it does not prove negative subscription gross margins. Anthropic’s API margin is estimated above 90%, while most subscribers leave the overwhelming majority of their allowance unused.

  • At roughly 10% utilization, Max’s tokenomics model puts Anthropic subscription gross margins around 50–60%, versus more than 90% for API sales. His provocative alternative explanation is that API tokens carry “punitive, exploitative, speculative pricing”; Jordan’s rejoinder was simpler: willing buyers are paying the market price.

  • The team rejects viral analyses assigning one universal API-equivalent value to a plan. Each plan is effectively a bundle of internal credits, and every model-token combination consumes those credits differently: “The $200 Anthropic plan” is worth a specific amount only for a specified model and workload.

2. Four token types turn one subscription into many different products

  • Andrew’s taxonomy starts with fresh input at roughly 1x cost, followed by cache writes near 1.25x and cache reads around 0.1x or less. Output—the model’s decoding work—is usually dearest: the example was $50 per million output tokens against $10 per million input tokens.

  • Cache reads retrieve an existing KV cache rather than reprocessing the whole conversation. Since an autoregressive chat repeatedly carries its history forward, cached context can dominate total volume; when most of the workload is cached, the cache-read price can almost completely determine the cost.

  • The physical analogy matters: cache reading is largely memory movement, while output creation is sequential decoding on the accelerator. That makes workload mix—not merely total token count—the load-bearing variable in any comparison.

3. The limit tests isolated every token price experimentally

  • Andrew’s team bought each plan and imitated the calls made by popular coding CLIs. For input and cache experiments, they placed a small tag before a token-heavy passage from War and Peace, then requested a trivial answer of roughly four tokens so output would not distort the result.

  • Randomizing the tag and omitting the cache header isolated fresh input; setting the header produced a cache write; repeating the identical cached request isolated reads. For output, they demanded a very long technical essay—up to 16,000 generated tokens—and measured how much of the weekly allowance disappeared.

  • Repetition produced curves from which the internal weights could be solved. Their observed agent workload was about 96% cache reads; the illustrative chat mix was stated, in rounded terms, as roughly 75% cached input, 10% uncached input, and 13% output tokens.

4. Cache-hit rate is an infrastructure moat hiding inside token prices

  • A production provider must preserve conversational state while serving thousands, tens of thousands, or millions of intervening queries: write tokens somewhere, keep hot state in HBM, demote it after a pause, then return the next request to the correct model deployment without leaving processing units overloaded or idle.

  • Time-to-live policy creates a direct user penalty. A cache might remain at the highest tier for five minutes or perhaps an hour; once an old conversation expires, restoring it through a full input or cache write can cost ten times as much—or more—than reading the live cache.

  • Max says early Anthropic ratios were roughly 1 for cache reads, 10 for input, and 50 for output, implying that Anthropic may have captured about 99% of the profit from cache reads. Competitive cuts have therefore targeted the highest-margin, highest-impact component: Astra still charged $1 per million cache-read tokens against $0.25 for Fable.

  • OpenRouter’s “effective price” is useful precisely because it incorporates an assumed cache-hit mix. The episode’s operational conclusion: high hit rates are not accounting trivia but difficult routing and memory engineering.

5. Anthropic wins the measured plans, though model quality remains disputed

  • On the $200 tiers, the tests yielded about 1.9 billion Astra tokens versus 3.1 billion Fable 5.1 tokens. Fable was also about 40% cheaper in API pricing in this comparison because its cache reads were 75% cheaper. Fable used only half of Anthropic’s allowance, leaving another 50% available for billions of Opus tokens—so the headline dollar chart understated Anthropic’s advantage.

  • Critics invoked Artificial Analysis figures of approximately $5 for a maximum-cost Opus 5.5 task versus $0.72 for a GPT-5.1-Codex task. Andrew’s pushback: the benchmark is imperfect, Opus 5.5 at medium roughly matches 6.1 Soul at X high, and that matched-quality cost gap is closer to 2–3x; against about 5x subscription value, Opus still delivers roughly twice the value.

  • The disagreement remains visible. Andrew calls Opus 5.5 a class above 6.1 Soul, while Max and Jordan stress that model preference is task-specific and personal. Andrew’s SemiAnalysis repository also automatically rejects pull requests unless the specified model is Claude Opus 4.5 or Claude Sonnet 4.5.

6. Model ownership, utilization, and task quality decide the durable economics

  • Moonshot appeared weakest among the cited Chinese plans at roughly 7x API-equivalent value, while some peers landed around 12–15x, near OpenAI. None offered such overwhelming value that users should abandon OpenAI or Anthropic for a weaker model; Meta’s $50 plan was highlighted as yielding about $2,500 of Spark usage, potentially 13 billion monthly tokens.

  • Max believes Anthropic utilization cannot be materially above 10% if leaked financials are to reconcile, leaving subscription margins around 40–70%. Power users with multiple accounts dominate online discussion, but corporate seats that consume only a fraction of their apparent API value can still satisfy both buyer and vendor.

  • OpenAI’s 50% plan-limit cut was announced around 9 p.m. on the eve of DevDay, surrounded by positive releases and softened by a one-month grace period. The speakers expect the response may change once grandfathered users feel the lower caps; they also emphasize that every estimate is point-in-time and could shift sharply with new models or pricing.

  • Token efficiency itself resists benchmarking. Andrew says a model such as Astra may complete the immediate task faster, while Opus or Fable may perform refinements immediately, producing a cleaner long-run codebase; the leaderboard cannot cleanly separate model quality from operator skill or quantify “how much of anything you do.”

  • The open-versus-frontier split ultimately concerns task demand. Simpler work can rationally move to cheaper open models, while frontier economics require new high-intelligence opportunities to grow faster; Andrew’s thought experiment asks whether the economy could absorb “100 million super-smart PhDs” at high returns.

  • Jordan’s cybersecurity experience shows both sides: refusals pushed him to GLM-5.3, whose existing context then handled an ordinary dashboard PR with far fewer tokens. Jordan’s conversations with well-known hackers suggested unlocked 5.6 Soul, Astra 6 Cyber, or Mythos 5.1 could be far stronger; Andrew separately said frontier-lab researchers may be genuinely concerned about cyber capabilities, while the exchange also raised—and muddied—the “regulatory grab” interpretation.

Full transcript
Jordan Nanos

Hello everyone, welcome back to SemiAnalysis Weekly. This week, I'm joined by Max and Andrew, authors of a very controversial but interesting article, “Anthropic Subscriptions Are Five Times More Profitable Than OpenAI.”

We published this material a few days ago, where the guys conducted limit testing of each AI tariff plan. It's not just OpenAI with Codex and Anthropic with Claude. This also includes Meta, SpaceX AI, MiniMax, Moonshot, Z.ai, Cursor, and Cognition—any service where you can purchase a comprehensive subscription plan.

We'll talk about the conclusions they drew. Guys, welcome to the show.

Max Kan

Thanks for the invitation, Jordan. First of all, a TV or a monitor or something like that—this studio is too professional now. I'm not used to this.

1. Subsidies and Credits

Jordan Nanos

Well, yes. Other people need me to guide them with questions, but for you, I think I can let you lead yourselves. Max, these plans are heavily subsidized, aren't they?

Max Kan

It depends on how you want to define subsidization. Compared to API prices, these plans are obviously more cost-effective. We can put a graphic on the screen, or people can read the news if they haven't seen it yet, but some of the results show that you can pay $200 for the Claude plan and get $12,000 worth of tokens at API prices. So this is definitely a good deal.

However, are these subscriptions subsidized in absolute terms, or does Anthropic have a negative gross margin on subscription sales? I think the answer is no. The reason is that, firstly, their margin on the API is simply incredible. The consensus forecast is over 90% margin at this point.

Secondly, most people don't use the service 100%. If you combine a realistic usage of, say, 10%, with that 90% margin on the API, it turns out that subscriptions are about 50%–60% of the gross margin for Anthropic. So it's still a pretty good business, but compared to their API business, with a margin of over 90%, subscriptions look so-so.

Jordan Nanos

Maybe the conclusion is not that subscription plans are subsidized, but that API pricing for current models is punitive, exploitative, speculative pricing—whatever you want to call it. It's pretty expensive.

Max Kan

They sell to willing buyers at a fair market price. This is capitalism.

Jordan Nanos

Yes, of course. Good. So tell me about 100% utilization and how you did the testing to get those numbers and results. This is a question for Andrew.

Andrew Megalaa

Yes. To figure out the results for subscription plans, we purchased each individual plan and measured each token type. Because of that, we can tell what type of workload we're running. In this case, it's generally agentless. That's exactly what I used these subscription plans for.

That's usually about 96% cache reads, a little bit of input—about 2% or so—and a little bit of output. Mix it together, and you get those big numbers we found.

Max Kan

Before you get into the details, I want to point out that there were already people on Twitter doing similar analyses. They were essentially just presenting a single dollar value as the equivalent of the API cost for a subscription. This is, in fact, an overly simplified and incorrect analysis.

The way to think about your subscription plan is that, by paying the company, say, $200 per month, you get a certain number of credits. It's just some other currency. You should think of each combination of model and token type as costing a different number of credits and consuming your limit differently.

Because the cost ratio between different token-type models may differ from their API cost ratio, the actual API cost equivalent of your subscription plan can vary dramatically depending on which model you use and what workload you are running. This means you can't just make a blanket statement like, “The $200 Anthropic plan gives you $10 of value.”

You should say, “The $200 Anthropic plan, when you do a certain workload on a certain model, gives you such-and-such a dollar amount.” This is the nuance that we highlight in our newsletter, and I really want to make sure people understand that.

2. Token Types and Cache

Jordan Nanos

Yes, that makes sense. Can you go back to basics and simply describe the different types of tokens? Since you mentioned token types, I'm not sure everyone understands what that means.

Andrew Megalaa

When you use the Claude Code CLI or the Codex application, there are different types of tokens used when communicating with the model. The first type is input tokens. Input tokens are not cached. These are simply new tokens that you send to the model.

When you have an agent-based, multipass dialogue, most of your input tokens come not from direct chat, but from, for example, web search, where you get something and then analyze it with a small subagent in one step. In this case, caching is not required, so the cache is not used, and this is considered the cost of the new input token.

The second type is a cache write. Caching means that during a multipass dialogue, instead of reprocessing the entire query each time you access the LLM API server, you can save previous results in a cache and reuse them for faster output and lower costs.

Cache-write tokens are previous tokens that you write to the cache. They're charged once on write, and then when you read them, they're charged as a cache read. The nice thing about reading from the cache is that it's extremely cheap, as you're effectively just getting the data from the KV cache on the LLM server.

So now you have input, cache write, and cache read. Inputs usually have a base price of 1×. A cache write typically costs about 1.25× the cost of the input data. Reading from cache costs about 0.1× or less of the cost of the input data.

Finally, you have the output tokens. This is the output data that the model directly generates. This is usually the most expensive type of token—for example, $50 per million tokens compared with $10 for input tokens.

Max Kan

Yes, it's quite simple. You can draw a parallel between the token type and the actual processing that happens on the accelerator that creates those tokens. In other words, reading the KV cache from memory is a cache read. This is significantly cheaper in terms of performance than creating new tokens, which is the decoding step, right?

Andrew Megalaa

Yes, that's right. And read caching is probably the most important thing in how you configure caching with an LLM, because these models work like autoregressive models. If you don't use caching, you process the entire request again each time, which is extremely time-consuming and expensive.

Most of your chat is actually cached reading because you're processing everything from scratch every time. If everything is cached, it's essentially a continuous cache read. You want the cache-read price to be as low as possible, because it almost completely determines the cost.

Jordan Nanos

Got it. Can you tell us about the technical side?

Max Kan

I was going to say that this was a bit of a tease for our internal X-endpoint test. We really have a high percentage of hits in the cache. This is exactly what distinguishes good inference providers from worthless ones. It's actually a surprisingly difficult problem, but it makes a huge difference in the cost of running the same workload, considering everything I just said about different token prices.

Jordan Nanos

Yes, and OpenRouter recently added this too. If you go to OpenRouter for any model and scroll down past the main provider prices, they have a great “effective price” that shows the real value of the mixed tokens, assuming a certain percentage of cache hits from different providers. That's pretty cool.

So explain why cache-hit rate, or CHR, is such an important metric for those doing inference, and why it's so difficult.

Andrew Megalaa

As we said, cache is a conversational state, right? Why is it difficult to direct a new request to where that history is stored, or to move it to a slower but cheaper storage tier?

If you're running a real production workload where there are many different racks and deployments of the same model, and each of them is processing a batch with a bunch of user requests, when I send, say, the first part of my request, you have to be smart enough to write those tokens to a cache somewhere.

Most likely, they'll stay in HBM at first if you're sure I will write the next prompt immediately. But if there's a big pause, you'll probably offload them to a cheaper memory tier. There's a circle of life there.

When I come back with the next step of my query, you've obviously already processed thousands, tens of thousands, or millions of other queries in the meantime. You have to take my new query, probably send it to the same model deployment, find my previous tokens—which you hopefully have cached somewhere—and then route them to the same deployment and process the next step.

This process of looking up the previous cache, routing the next steps of the same request to the same deployment, and balancing everything properly so that none of your PPUs are overloaded, while avoiding a bunch of PPUs sitting idle and burning money, is all very complex systems design.

Max Kan

Yes, and there are TTL dynamics here too, right? Each new request is assigned a lifetime. It might stay in HBM, at the highest level, for maybe 5 minutes, and another one might stay for an hour or something.

But it's a penalty for the user if you let your old conversation die. When it's no longer cached and you want to continue, you have to pay the full cost of the input, or a cache write that's 10 times more expensive, to restore it. Often, at this point, it's even more than 10 times.

One of the things we pointed out to our All Access subscribers a few months ago is that, if you look at the initial pricing ratios of the Anthropic API, if the cache read was 1, the input tokens were about 10 and the output tokens were about 50. It was a ratio of 1 to 10 to 50.

We noticed recently that, with the price cuts from OpenAI and Anthropic, they were mostly focused on making cache reads cheaper. Why is this happening? First of all, as we have already found out, this is the most important lever for a real reduction in the total effective price. Secondly, as we pointed out to our All Access subscribers a few months ago, cache reads were initially by far the highest-margin token type of the three. It is likely that, at the initial price ratio, Anthropic received about 99% of the profit from cache reads, which is simply incredible.

Now we are finally seeing these prices come down again under competitive pressure. This is one of the points I want to draw attention to. Unfortunately, I don't think this has reached OpenAI yet. Astra still costs $1 per cache read, compared to 25 cents for Fable. Yes, yes, yes. This is actually a good reason why Astra may have a higher dollar value in the Pro 200 plan compared to Fable 5.1 in the pricing plans. Although in reality Fable 5.1 has more tokens and more usage.

Because the cost of reading from the cache is 25 cents per million tokens versus $1 in Astra. And that's not to mention that Fable 5.1 only takes up 50% of your plan, and you have another 50% to use.

3. Model Value Charts

Jordan Nanos

So, explain to me in detail how OpenAI and Anthropic compare in terms of their leading models in these plans.

Max Kan

This is the main graph that many of those reading the article have seen. It shows how Astra directly compare to Fable 5.1 across the different plans they offer, in terms of how much value you get per token when using these models.

Choosing which model you prefer now seems like a very personal matter. I don't think anyone could say, at least for my tasks, that Astra is better than Fable or Fable is better than Astra. They just seem to offer different benefits depending on the tasks I'm trying to accomplish.

Jordan Nanos

But tell me about the actual usage.

Max Kan

Yeah. Strangely enough, I think this is not the graph that gained popularity on Twitter. We can discuss that next. People didn't criticize this one because it shows that OpenAI and Anthropic are roughly comparable for the same $200 or $100 a month, or whatever it is.

It's worth emphasizing one point that people missed and that Andrew talked about earlier. Initially, it was claimed that our different schedule was supposedly unfair to GPT 6.1 Soul because the API price is lower than Opus.

Regarding this main graph, where we compare Astra and Fable: Fable is actually 40% cheaper in API pricing than Astra because their cache reads are 75% cheaper, which is over 9% of your total tokens. If you look at it from a token-count perspective, on the $200 per month plan you get 1.9 billion Astra tokens, but on the same $200 plan you get about 3.1 billion Fable tokens.

Not only that, but those 3.1 billion tokens represent only 50% of your limit, not the full 100%, as Fable usage was limited to only 50% of your token limit. So you can actually use billions more Opus tokens on top of that for the same $200. I think even this graph clearly shows that Anthropic offers better value.

If we were to go to a Soul vs. Opus chart, I think it would show an even more crushing dominance by Anthropic. Maybe Andrew can discuss this.

Andrew Megalaa

Yeah, that was a very controversial graph on Twitter, or X, where everyone was saying, “This is so unfair because the 6.1 Soul is so much more efficient per token than the Opus 5.5.”

Everyone kept citing the Artificial Analysis Intelligence Index analysis of the maximum cost of Opus 5.5 per task and the maximum cost of 6.1 Soul per task, which came out to something like $5 versus 72 cents for GPT-5.1-Codex. People used that to justify what they called an unfair comparison.

There are several reasons why this is problematic. Even assuming that GPT-5.1-Codex is indeed that much more efficient in terms of tokens, mathematically it still works out that Claude Opus 4.5 provides double or triple the value in terms of the number of tokens.

The exact math is this: first, we disagree that the Artificial Analysis Intelligence Index is necessarily a good way to measure the effectiveness of token usage. We will return to this issue later. Even assuming that using these artificial, often imperfect and oversaturated benchmarks is a good way to measure token performance, the Opus 5.5 at max is the number one model in the index, while the 6.1 Soul is not.

And if you just use the Opus 5.5 on medium settings, it beats all of the 6.1 Soul's results, except perhaps max. So, I think specifically the 6.1 Soul and X high, as well as the Opus 5.5 on the medium, have the same scores on the AI Intelligence Index. If you look at the cost per task for those 2 specific runs, the difference is about 2 to 3 times. Then you get 5 times the equivalent API value according to this table.

In reality, you still get twice as much value from Claude Opus 4.5 as from GPT-5.1-Codex, even if you assume that this imperfect benchmark index is a good indicator of token performance.

Max Kan

Yes, it's a shame that the Opus 5.5 doesn't really belong in the 6.1 Soul class. It actually competes directly with Astra.

Andrew Megalaa

I don't know if I completely agree with that.

Max Kan

A lot of people are replacing, for example, GPT-5.1-Codex-Max with Claude Sonnet 4.5. This is Andrew's point of view as an Anthropic fan, okay?

Andrew Megalaa

Maybe that's my point of view. This could be my point of view. But I would say that Opus 5.5 even on imperfect benchmarks still remains number one in a good intelligence index.

A lot of people are replacing, for example, GPT-5.1-Codex-Max with Claude Sonnet 4.5. Anthropic says it is as good as, if not better than, GPT-5.1-Codex-Max. So, in that scenario, you have Claude Sonnet 4.5 as a competitor to GPT-5.1-Codex-Max at half the cost of your plan. That's essentially equivalent to the $200 plan you already get on ChatGPT.

But you also have Claude Opus 4.5, which is a beast for the rest of your plan, costing around 5 or 6 times as much. When you compare that the question of who provides more value compared to the H 2 BT doesn't even stand up. H 2 BT vs. Anthropic subscriptions. Because Opus 5.5 is simply a level above 6.1 in most categories.

Max Kan

For the audience, Andrew has a repository on SemiAnalysis. he has an agent.md where you have to specify which model was used to create a pull request, and if it's not Claude Opus 4.5 or Claude Sonnet 4.5, it automatically rejects the pull request.

Andrew Megalaa

Yes. He's actually not even allowed to contribute to this repository.

Max Kan

Yes, okay. That's pretty much what I was saying earlier. At this point, choosing a model is a very personal decision, and it's not necessarily for everyone. Andrew is forcing all GPT-5.1-Codex-Max fans to upgrade to Claude Opus 4.5 if they want to contribute to this repository, and they saw the light. Everyone who started doing this converted.

Andrew Megalaa

This is true. This is true. It works. They saw the light.

4. Test Methodology

Jordan Nanos

Can you tell me a little more about the methodology? How did you actually measure the number of tokens of each type and how the limits are filled? I think that's what confuses people, too. They run one really big request, and their hourly limit, 5-hour limit, or whatever it is just runs out.

You mentioned at the beginning, Max, that no one uses these things 100% of the time. What scenario did you actually use to figure out how many tokens could be consumed in these time periods?

Max Kan

Instead of running some task or test and seeing how many tokens it uses, we wanted to conduct a kind of scientific experiment: extract the exact values for input tokens, cache writes, cache reads, and output token prices. We do this by mimicking how all popular command-line interfaces call the model API. We're simulating how you would access the inference server as a regular user with a subscription.

We use 2 types of queries that allow us to maximize the type of tokens we're interested in and minimize the types we're not interested in. For input, cache writes, and cache reads, we use a small tag at the beginning of the request, followed by a whole block of text from War and Peace that takes up a large number of tokens. Then we ask the LLM a very simple question to minimize the number of output tokens it generates in response. It's usually about 4 tokens that just point to history, fiction, or something like that.

The reason this works is that, for input tokens, we can set a cache header to indicate whether something needs to be cached or not. We can also randomize the tag before the initial prompt, which essentially ensures that nothing gets cached, and that the API doesn't try to cache it since there is no cache header. This is present in the Anthropic API, OpenAI, and other models.

Then we read the response to check what types of tokens were actually used, so we know what was consumed. For caching, the situation is similar to input, except that we mark it for caching, so we can see that the entire prompt is cached, except perhaps for the last model-generated responses or output tokens.

For reading from the cache, we run a similar experiment and simply set the cache header. One call will be a cache write, which we drop from our count. We then continue sending the same request, and that gives us a cache read for each subsequent request.

Finally, for output tokens, we write a very, very long technical essay, and then the model generates 16,000 output tokens. This consumes our entire budget, and we can look at the response and the counters and see that it used about 2% or 5% of our weekly limit. Output tokens are expensive. They probably use 5% to 10% of the limit in a single call.

By repeating this and building graphs, you can do the calculations and determine exactly how much each type of token is worth. We can then use this for any analysis. For example, if you're running a light load, then with this token distribution, this will be the cost of your subscription.

And with a certain chat workflow, it will be the cost of your subscription, and so on. It’s also worth discussing how you get these ratios. We keep a good eye on what the chat workload actually looks like compared with the agent workload, right—in terms of the input-to-output ratio, cache reads, and cache writes.

Jordan Nanos

Mm-hmm.

Max Kan

Yes, we took the data on agents from our economic panel, where we have the internal SemiAnalysis statistics for all months, and we used the data for September. It was one of the freshest months, so this was our own token distribution.

Everything we do in corporate accounts is the work of agents. Chats are a bit more complicated, because the only way to get a real token distribution is to have ChatGPT with a real workload. But this can be roughly estimated based on the fact that there is a system prompt for each message in the chat, or for each chat thread.

That system prompt is essentially cached for all users of Claude, ChatGPT, and other services. So this alone already accounts for a significant portion of cached reads. That’s where we start, and then we assume that users send fairly short queries. Maybe they leave and don’t respond to the thread anymore, or they come back and write once or twice.

Under these assumptions, we get the distribution given in the article: somewhere around 5% of the input data and approximately 20% of the cached reads. Yes, 13% output tokens, 75% cached input tokens, and 10% of the uncached input tokens. Logically.

5. Other Plans

Jordan Nanos

Makes sense. What about other coding plans? For example, one of the subheadings talks about other consumer plans that aren’t from Anthropic or OpenAI. First of all, does this even matter to anyone? But let’s assume that it does.

What do you think of Muse by Meta, MiniMax, Super Grok Heavy, Cursor Composer, GLM, and Kimmy, and how do they rate their value compared to others? I was a bit surprised that some Chinese developers, like Moonshot, Z.ai, and MiniMax, offer pretty good API value in dollar terms.

Max Kan

The reason this surprised me is that they all have low computational costs. I think if you look at their API prices, the margins there are even lower than those of the leading labs, because it’s a more competitive business. It’s more like a commodity market.

Some of them—Moonshot seems to be the worst—offer about 7 times the value for the equivalent API. But some of them were in the same range, at 12 to 15 times the value, as OpenAI. Considering that OpenAI is the complete opposite in terms of compute costs and has cheap API prices, at least for GPT-5, this was a rather unexpected conclusion for me.

Jordan Nanos

But that’s not true. I guess my main takeaway is that when you look at these graphs, it isn’t obvious that any of these pricing tiers are so much better value for money that you’d want to walk away from OpenAI or Claude, thinking, “Okay, I’m getting a much better return on my Manus subscription, so I’m willing to use a slightly worse model because I get a lot more tokens.”

This is probably 2 or 3 times higher than the OpenAI or Anthropic limits. Or maybe it’s a good thing that it starts at $50 instead of $200 or $500 for access to their top-of-the-line model. But as far as we can see, no one is having a price war on monthly subscriptions right now, right?

Max Kan

Yes, that’s true. I think the main reason is that no one can match Anthropic’s margin on API prices. So if they offer similar equivalent value per dollar, that probably means their real margin is much worse than Anthropic’s.

This only highlights that, in the age of artificial intelligence, having a model that everyone wants to use is a huge advantage. Being a wrapper is especially disadvantageous.

We covered this at the end of the article, but if you look at the use of Anthropic models in Perplexity or Cognition, the limits there are much lower than at the original source. This is obvious if you think about it for 2 seconds, because Anthropic has over 90% of the gross profit, while services like Perplexity or Cognition have, at best, about 20% of the API cost or something like that, compared with API prices.

But it was pretty cool to see that this claim was backed up by data.

Jordan Nanos

Yes, that makes sense. How about a little thank you to Meta for one of the best value plans out there. With a $50 Meta subscription, you can get about $2,500 worth of new Spark features. That’s probably somewhere between 140 million and 260 million tokens per dollar, or 13 billion tokens per month. That’s quite a lot.

If they keep the same rate and the new Manus is a little better, it could be a great deal.

Max Kan

Yes, dude. For just this $50 plan, you can get around 250, maybe 300, tennis-court reservations in San Francisco through the new Spark.

Court reservations for a dollar?

6. Utilization and Limits

Jordan Nanos

Yeah, I’m not sure, dude. I think—yes, yes. Okay, let me ask you a question about this. We’ve posted a similar chart before, Max, and I found it extremely interesting then, and it’s even more interesting now.

We don’t see what the real usage rates of these plans are. Of course, there are some loud “pro users” on X who use all 7 plans at 100%, squeezing the most out of them. And there are other people, a little more sensible, like me, who spend half the day recording podcasts, don’t use my coding agent, or let it run in the background until it freezes, and then I have to move on.

I’m dozing, sleeping, you know.

Max Kan

Are you napping, seriously?

Jordan Nanos

I’m not napping, but—never mind. Yes, yes, but wait. Of course, yes, yes. My daughter is napping.

So, if you had to guess, what percentage gross margin do these companies actually have on the coding plans for these models? Where would you place them? Does the average user use 10%, 100%, or something in between?

Max Kan

I think for Anthropic, this figure cannot be much higher than 10%. The reason is that if you look closely at all their financial-data leaks, the numbers don’t add up if the subscription business has extremely negative margins. I think it has to be a business with a decent margin—say, 40% to 70%—for all the financial-statement leaks to converge.

We have broken down all the details in our tokenomics model. That’s why I believe the utilization rate cannot be much higher than 10%. This also makes sense, considering that many subscribers are actually corporate customers or use business plans.

If you’re a business buying a $1,000-a-month plan for your employees, and they're only spending $800 or $1,000 a month on credit, you’re still happy with that deal. That’s a utilization rate of less than 10%.

Jordan Nanos

Right. In your discussions with people who have these plans, what is their first reaction to what you’re saying? It seems that at first, anyone who has tried paying for tokens and then had access to unlimited plans says something like, “I will never pay for tokens voluntarily, because these unlimited plans are a bargain.”

Roughly speaking, what is the reaction when you talk to them about the article? Do they agree or disagree, or do you think people are at 10% to 20% usage or thereabouts?

Max Kan

The guys on Twitter who wrote a lot about it definitely didn’t agree. These are the 100% users, with 5 accounts, creating new ones every 3 days or something like that.

But I think their sense of the plans somewhat coincides with our actual data. Many people have been saying in recent weeks that Opus 5.5 Auto Cloud 1 plan is, in particular, the best value for money you can get.

It’s also important to note that the 50% reduction in limits for the $200-per-month plans that OpenAI announced at DevDay has not actually taken effect yet for those who purchased the plan before the announcement. They received a 1-month grace period from OpenAI.

So we were pretty surprised that Twitter didn’t smash OpenAI harder for this 50% limit cut, and for providing a bunch of code saying, “Oh, since GPT-7 Luna will be smarter than GPT- 6 Astra, that’s a win for you.” It was kind of pointless, but no one criticized them for it.

I think this is largely because they timed it well with a bunch of positive PR on DevDay, and because they provided that month-long grace period. Maybe in a month, when the limits start to really affect most users, we’ll see more buzz around it.

Jordan Nanos

It was good of them to announce the cuts around 9 p.m. on the eve of DevDay, and then cover them the next day with 20 awesome announcements about all sorts of other things.

Max Kan

Dude, just imagine if Anthropic announced this. They would simply be eaten alive. They would become a meme.

7. Token Efficiency

Jordan Nanos

Logically. Um, it's a good thing we haven't discussed what's worth talking about here yet. Virtually everything related to subscription restrictions. It's probably worth emphasizing that this mailing material is very much tied to a point in time. As new models come out, new subscription levels appear, and new things like ultrafast beta become available—or providers decide, “Hey, I actually want an 8% margin on my subscriptions, not 50%”—we expect all of these numbers to change quite dramatically.

We will constantly monitor this and let you know about updates as they come out. We’ll also keep track of model quality and token efficiency, which will definitely come in handy, I’m sure.

Max Kan

Dude, it’s so hard to actually measure token efficiency well. If only there were some way to do it.

Jordan Nanos

Yes. You can’t rely on benchmark tests, man.

Max Kan

But without benchmark tests, how can you actually verify that 2 models have the same quality?

Jordan Nanos

I don’t know, man. Give 2 people the same task, check in a week, evaluate the quality of the work, and see how many tokens they used.

Max Kan

I think there’s a huge hidden variable here, which is the person doing the task.

Jordan Nanos

So what do you think has the greatest impact on token efficiency: 1 person using 2 different models, say Astro and Fable, or 2 different people using the same model to achieve a goal? Andrew has a strong opinion on this matter. He came, looked at the SemiAnalysis token-efficiency leaderboard, and developed his own strong opinion on this.

Andrew Megalaa

So, if you look at the leaderboard, people who are top users on Astra and Fable are spending about the same amount, but Fable users are using more tokens. That doesn't really give us much information about how much of anything you do, which is the hardest part of answering this question. I would say from personal experience trying to work with 3 models, the effectiveness of OpenAI tokens is definitely noticeable.

I think of it as short-term token efficiency. Yes, I'll complete this immediate task, but I feel like there are so many refinements I'd like to add that Opus or Fable would do right away, while Sonnet or Astra could ignore them to complete the task faster. In that sense, in the long run, I feel that better models create a better codebase that is much more readable.

If you read some of these Astra tests, that's something. But that's just my opinion. I think the models themselves have a pretty big impact on the performance of tokens, even more than which user uses them.

Max Kan

What about resellers of these tools, like Cursor, Cognition, and Perplexity? How do you assess their influence on the choice of model that you can use? Have you ever considered using a router where you switch models, or letting Perplexity or someone else choose a model for a specific task because they think it's better, even though their choice is different from yours?

Andrew Megalaa

I don't use Cursor or Devin myself, but if I did, I'd probably still choose my favorite model. I know Devin has Suite 2 and Fusion, which supposedly use a cheaper model along with a bigger and better one to keep prices low. It looks great on benchmarks, but I haven't personally tested it in real-world conditions.

My overall impression is that if I can afford it and my company is willing to pay, I would definitely choose Opus over a cheaper model like the Suite 2, or even the cheaper Sonic 55, simply because I trust these models to always get things right. I don't have the same confidence in other, lesser-known models.

I think when you work with them, you start to understand what each model is good at and where it can let you down. That is actually quite important when designing systems.

8. Open Models

Max Kan

What about the growing popularity of open models? I saw David Friedberg's statement on the All-In podcast that there has been a shift from closed to open models over the past 12 weeks, from an 80/20 to a 20/80 ratio. We wrote a whole article for our Tokenomics subscribers about how that's not entirely true, because they're only looking at API data.

But I realized that there's a pretty significant shift happening, at least in terms of token volume—though maybe not quite in terms of revenue—where people are really using open models more often. Do you have any personal experience or conversations with others that lead you to believe that frontier labs are doing something wrong with the pricing or quality of new models? Are they not sufficiently ahead of the open ecosystem, forcing people to switch to open models and stay with them?

Andrew Megalaa

I guess the question is how far ahead the closed-source models are. Actually, I think it's just a matter of how high-quality a model I need for my task. It's true that many fields of software development, and intellectual work in general, don't require the intelligence of GPT-4 or GPT-5. Therefore, I believe that many businesses, especially those with low margins, are making a rational decision to delegate simpler tasks to these increasingly powerful open models.

I think this will continue to happen in the future. The question is whether the frontier-lab business model is sustainable. It comes down to your belief that the opportunity space for even smarter AI will outweigh the share of tasks moving to cheaper models.

I think this is perhaps the biggest difference in views between those who believe in frontier labs and those who believe in open source. If you magically created 100 million super-smart PhDs—experts in all fields who didn't need to eat or sleep—would the economy be able to absorb that quickly with a high rate of return?

Do you think that, on a planetary scale, the world is simply not smart enough to absorb this? I think sensible people have very different views on this issue. If you believe the answer to that question is yes, you're optimistic about Anthropic. If the answer is no, you probably think that we're going into a recession relatively soon. If you're really optimistic, type together.

9. Cyber and Capture

Jordan Nanos

Yeah, interesting. I have a reason to ask this question because I've been using Perplexity so much in Slack lately, specifically GLM-5.3, because of all the denials from OpenAI and Anthropic on anything cybersecurity-related, which made my job impossible. I had to use GLM-5.3, and I think that led me to discover that it's pretty good for a lot of things.

They have a model with a level of harm, right? I mean, it was just amazing in cybersecurity tests, and I literally can't test it. For me, this is the advanced cybersecurity model because I literally can't use any other model. It doesn't matter—I can only look at the model scorecards. I can't really use other things, even though I'm in the CVP program to get verified for all this cyber work.

I still can't get around these restrictions. I need a more creative approach to prompts to get around the limitations built into the model itself, rather than the classifier outside of it. But anyway, it's a waste of time trying to convince a model that my grandmother is in danger and that I'll bribe her with a peanut butter sandwich if she doesn't do this thing.

It's easier to switch to GLM, and it will do the job. What happened is that I was in the same chat history, with all the data there, and I didn't switch to something else. I just asked it to make a PR for the dashboard repository to post the data we got as a result of that work, using 10 times fewer tokens.

Max Kan

As for the price, she did this PR for the dashboard. Hmm, no problem. I didn't really need Fable 5.1. For, you know, 10 lines of code. This is also music to Dylan Patel's ears. Many things definitely don't require GPT-5.1.

Jordan Nanos

Yes, but I also feel like I use these things day and night. My wife will tell you that I stay up late using them, and I don't make the top 10 on the leaderboard. So I still have no idea what the hell you actually do with all these tokens.

Max Kan

We're constantly out of the top 10, you guys. You really should interrogate guys like Jeremy, Kyle, Andrew Wagner, and the others. Jeremy will be back soon. When I'm done asking him about data centers, I'll ask him what this fast mode actually gives him and how many subagents he uses.

He told me he was just taking it. He once admitted that he chooses arbitrary odd numbers and says that this is exactly how many subagents to run for a task that doesn't require them at all. It just feels like he's saying, "Give me 37 agents to investigate all the permits for this one data center site."

It is a matter of honor for them to do this. Then he's like, "Oh my God. That was $2,000. Oh well."

What? This is research, guys. This is research. Don't touch my husband. Don't question my methods.

Jordan Nanos

You know, I want to add something else about GLM-5.3. After talking to very well-known hackers recently—really world-leading hackers—it seems that even the Soul 5.6, if it has the cyber capabilities unlocked, is a much, much better model than GLM-5.3 in real cyber tasks.

While some of these models may show high results in benchmarks, I think that Soul 5.6 Cyber, Astra 6 Cyber, and Mythos 5.1 with unlocked capabilities can be much more powerful than many people think.

Andrew Megalaa

Yes, there is something a little scary. There is a real possibility that the people working at these frontier labs have access to the models earlier than others and understand their capabilities better. They are not necessarily lying or building schemes for complete regulatory takeover, but are genuinely concerned about the cyber capabilities of these models.

Jordan Nanos

Do you think it's possible that this is exactly the scenario that's unfolding now?

Andrew Megalaa

I think it's quite possible. But the question is whether they are seeking regulatory takeover.

Jordan Nanos

No, I'm just giving a short answer. I mean, the researchers are sincere when they say—

Andrew Megalaa

I'm giving you a short answer to that question. Maybe researchers—

Jordan Nanos

Sorry, sorry. OpenAI just released a new repository called openai/math, where they seem to have made more progress in this field than in the last 10 years combined. I got distracted.

Andrew Megalaa

Sorry about that, but I didn't hear your question.

Jordan Nanos

I think they just solved the Millennium Prize problem.

Andrew Megalaa

I think that's a better answer to my question than what I asked, which was, "Yes, maybe the people who see these breakthroughs months before us and are really worried about what's coming next are being genuine in their concerns, rather than trying to implement some convoluted regulatory takeover scheme."

Jordan Nanos

That's exactly what I said.

Andrew Megalaa

More broadly and, well, this, this, this is a ridiculous idea, Jordan. This is clearly a regulatory grab. Oh my God. Anyway, guys, I think we've gone off topic enough, we need to read an OpenAI blog post about some crazy, uh, mathematical stuff. So it seems the entire repository on GitHub is openai/math. It's like a high level and just a bunch of stuff. We'll see you later. And then, and then, and then about 10 more tasks that I've never heard of, but which are probably very important. I will study math tonight. Well, guys, thanks for joining. Uh, we'll talk about subscription limits and token efficiency another time. Sounds good. Thank you for inviting us. Thank you, Jordan. See you guys. See you later.