Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status
Back to Originals

OpenAI Shipped a Product With No Price. Almost Every Headline Rate This Week Has an Expiry Date.

Kira Nolan··7 min read
Inference Economics · Pricing

On Thursday, August 13, OpenAI previewed a new API service tier called Ultrafast. It runs GPT-5.6 Sol, the same weights and the same intelligence as the Standard tier, at a claimed 750 output tokens per second, which OpenAI describes as up to 14 times faster than Standard processing. Cerebras is underneath it.

There is no price. There is no GA date. As of this writing there is not even a model ID string you could put in a config file. A small set of customers across coding, financial research, voice AI, and e-commerce got access, and OpenAI said access expands as capacity grows.

I have been staring at that missing number for two days and I have decided it is the most honest thing anyone published this week. Not because OpenAI meant it that way, but because it accidentally admits what the rest of the week's pricing news was working hard to disguise: the list price on a frontier API has stopped being a number you can plan against.

What Ultrafast Actually Is

The mechanism is not mysterious and it is not a model change. Cerebras builds a wafer-sized processor with roughly 44 GB of SRAM directly on chip, which means the weights for a served model can sit in on-chip memory rather than being pulled across a memory bus on every forward pass. GPU inference at frontier scale is bandwidth bound far more than it is compute bound. Remove the bus and the token rate moves in a way no amount of batching on conventional hardware reproduces.

This is the payoff on a deal signed in January. OpenAI contracted for 750 megawatts of Cerebras inference capacity delivered in stages through 2028. Reuters put it above $10 billion at signing; Cerebras later told investors the agreement exceeds $20 billion, and the structure includes warrants for roughly 10 percent of Cerebras plus about $1 billion in working capital from OpenAI to help build the sites OpenAI then rents. We wrote about the wafer-scale architecture bet back when it was still a thesis. Ultrafast is the first consumer-visible product to come out of it.

One number nobody published: the baseline. If Ultrafast tops out at 750 tokens per second and that is up to 14 times Standard, the implied Standard rate is somewhere around 54 tokens per second. Both figures carry an "up to" so dividing them is an inference rather than a measurement, and I want to flag that rather than launder it. But if it is roughly right, it tells you something uncomfortable about what you have been paying $30 per million output tokens for.

The Expiry Column Is the Article

Now put Ultrafast next to everything else that moved in the last ten days. I am adding a column most pricing tables leave off.

ModelHeadline (in / out per 1M)What Happens Next
GPT-5.6 Sol Ultrafastnot publishedLimited preview, no rate, no date
Gemini 3.7 Flash$0.75 / $3.75Doubles to $1.50 / $7.50 on Jan 1, 2027
Grok 4.6$2.00 / $6.00Whole request rebills at $4 / $12 past 200K tokens
Claude Sonnet 5intro rateReverts to $3 / $15 on Aug 31, 2026
DeepSeek V4 Pro$0.435 / $0.87Increase announced Aug 6, amount and date unspecified
GPT-5.6 Sol Standard$5.00 / $30.00Unchanged

One stable rate in the group, and it belongs to the most expensive model on the list. Everything else is either promotional, threshold-triggered, expiring on a calendar date, or unannounced. Gemini 3.7 Flash is genuinely half the price of its three-week-old predecessor, and it is half price for 138 days. Grok 4.6 held its rate at $2 and $6, and the moment your agent's context crosses 200K the entire request rebills at double, which for long-horizon agent work is not an edge case, it is Tuesday.

You can check any of this against live rates on our models tracker and run your own token mix through the cost calculator. Please do, because a spreadsheet built on any of the five rates above is a spreadsheet with a fuse in it.

The Axis Moved and Nobody Announced It

For three years the competitive question was dollars per million tokens. That question is quietly dying, and this week is where I would date the certificate. Here is the same set of models on the two axes that are replacing it.

ModelOutput speedAA Intelligence IndexSource
GPT-5.6 Sol Ultrafastup to 750 tok/s61Vendor / AA
Gemini 3.7 Flash~340 tok/s56AA
Grok 4.6not measured here61AA
Claude Opus 5not measured here63AA
Claude Fable 5not measured here62AA
GPT-5.6 Sol Standard~54 tok/s (implied)61Derived

Look at the top and bottom rows. Same model. Same intelligence score. A speed difference large enough that Artificial Analysis has started publishing time per task alongside capability, and Gemini 3.7 Flash's headline achievement is not a benchmark at all, it is sitting on the Pareto frontier of intelligence against wall-clock time at 1.7 minutes per task.

When two SKUs share weights and differ only in latency, you are no longer buying tokens. You are buying time, and time does not have a natural per-million unit. That is why the price is missing. It is not an oversight. It is a category problem OpenAI has not solved yet.

The Case That This Is Great, Made Properly

I want to give the optimistic read real weight, because I think it is largely correct.

Latency is not a nice-to-have, it is a capability gate. A voice agent at 54 tokens per second is a product where the human waits. At 750 it is a conversation. An agent that runs a 40-step plan at Standard speed is a background job you check on later; at 14x it is something you supervise interactively, and interactive supervision is the single best known mitigation for the agent failure modes this whole industry has spent the summer apologizing for. The four verticals OpenAI picked for the preview (coding, financial research, voice, e-commerce) are exactly the four where the latency cliff is a business model, not a preference.

And the economics genuinely do change shape. If a task completes in a tenth of the time, the operationally relevant metric is cost per completed task, not cost per token, and a more expensive per-token tier can win that comparison outright. Anyone still optimizing a dollars-per-million column is optimizing a proxy.

Where I Get Off

All of that is true and none of it makes an unpriced tier a product. It makes it a demand-discovery experiment, and I am fairly sure that is what it is. Pick a handful of customers in the four verticals with the highest willingness to pay for latency, watch what they do with it, then set a price against observed value rather than against cost. That is a completely rational way to run a launch. It is also the opposite of the thing a platform needs to be, which is predictable.

There is a supply-side reading too, and DeepSeek just illustrated it. DeepSeek matched OpenAI's 80 percent Luna cut within hours, took on something like 7 trillion tokens a week, discovered its fleet could not hold it, and on August 6 warned developers of a significant price increase it still has not quantified. Cheap tokens are a promise about capacity, and capacity is the part nobody can fake. "Access expands as capacity grows" is OpenAI saying the same thing more gracefully.

So the honest summary of the week is not that OpenAI got faster. It is that five vendors published five headline prices and four of them are contingent on a date, a token threshold, or an announcement that has not happened yet.

What I Would Actually Do

Stop modeling on list price. Model on cost per completed task, measured on your own traffic, with your own harness. That number already exists in your logs and it is the only one that survives a promo expiring.

Then put the expiry dates in your calendar rather than your memory. August 31 for Sonnet 5. January 1 for Gemini 3.7 Flash. Some unknown day for DeepSeek. If a rate change would break your margin, you need the alert before the invoice, not after.

Instrument latency as a first-class metric next to spend. If you are not logging tokens per second per provider today, you cannot evaluate a speed tier when it gets a price, and you will end up buying it on a press release. Provider reliability and incident history live on our status page if you want the availability half of that picture.

And do not architect a critical path around a preview tier with no model ID. That should go without saying. It will not.

Our Take

The interesting artifact of this week is a blank cell. OpenAI can build a wafer-scale inference stack and sell frontier intelligence at conversational speed, and it cannot yet tell you what that costs. Google can halve a price and only guarantee it for four months. xAI can hold a rate and double it at a context threshold most agent workloads cross routinely. DeepSeek can win on price and then lose on physics.

None of that is a scandal. It is what a market looks like when the unit of value is changing underneath the unit of billing. The vendors will figure out how to price time. Until they do, the number on the pricing page is marketing, and the number in your logs is the truth.

Three things I am watching. Whether Ultrafast ships with a per-token premium or a different billing shape entirely, because that choice defines the category. Whether anyone independently measures Standard Sol's baseline token rate and confirms or kills the 54 figure. And whether Anthropic or Google answers with a latency tier of their own before Q4, which would settle whether this is an OpenAI hardware quirk or the new second axis everyone has to compete on.