Google Shipped the First 2nm Phone and Took 4GB of RAM Out of the Pro. On-Device AI Is Still a Cloud Product.
Google held Made by Google 2026 on Wednesday, August 12, and the headline is legitimately a first. Tensor G6 is the first chip built on TSMC's 2nm process to reach a shipping phone, roughly a month ahead of whatever Apple puts in the iPhone 18 in September. The new TPU carries 50 percent more compute and twice the memory bandwidth of the G5. Google says on-device AI tasks run up to 3.5 times faster while using up to 3.5 times less energy.
Then read the spec sheet for the Pixel 11 Pro. Twelve gigabytes of RAM on the 256GB model. Last year's Pro shipped with sixteen. Google removed four gigabytes of memory from the phone whose entire pitch is on-device intelligence, and credited “software and silicon optimization” for the fact that it still feels fast. If you want the 16GB back, you buy the 512GB or 1TB configuration.
That is the story. Not the process node. Memory is the resource that decides how large a model can sit resident on a device, and Google just moved it in the wrong direction on its most expensive slab of glass.
The Lineup, and the Memory Line
| Device | Price | RAM at base | Note |
|---|---|---|---|
| Pixel 11 | $899 | 12 GB | 256GB is the new baseline |
| Pixel 11 Pro | $1,099 | 12 GB | Down 4GB from last year's Pro |
| Pixel 11 Pro XL | $1,299 | 12 GB | 16GB only at 512GB or 1TB |
| Pixel 11 Pro Fold | $1,899 | 16 GB | Up $100; the only 16GB base in the line |
| Pixel Watch 5 | $399 | 3 GB | Up from 2GB; first RAM bump in the line |
The Fold is the tell inside the tell. It is the one device in the family that kept 16GB at base, and it is also the one that took a $100 price increase. Google knows exactly what the memory is worth. It just decided most buyers would not price it.
Twelve Gigabytes Shared Versus Twenty Dedicated
Put this launch next to the one from two days earlier and the ceiling becomes obvious. On Monday, Meta published Muse Glimmer, a roughly 30 billion parameter agentic model that runs on a single consumer GPU, compressed to just under 20GB at about 4-bit. That is the current floor for a genuinely capable local agent: twenty gigabytes, dedicated, on a card that draws hundreds of watts.
The Pixel 11 Pro has twelve, and it is not dedicated. That pool holds the operating system, the browser, the camera pipeline that Google spent most of the keynote on, and every background app. The slice actually available to a resident language model is a fraction of it. Gemini Nano 4 fits in that slice. A 30B agent does not, and no process shrink changes that arithmetic, because 2nm buys you transistors and efficiency, not capacity.
This is why I keep saying the TOPS number on a phone chip is the least interesting figure on the page. Compute determines how fast you run the model you can hold. Memory determines which model that is. Google doubled the TPU's memory bandwidth, which is a real and meaningful win for the model it already ships. It did not raise the ceiling on what that model could be.
What Actually Stays Local
Sort the announced features by where the inference happens and the picture is clean.
| Feature | Where it runs |
|---|---|
| Camera Looks, Instant Night Sight, 4K Bokeh video | On device (ISP and custom accelerator) |
| Gboard Rambler speech cleanup | On device (Gemini Nano 4) |
| Live Translate in calls, podcasts, video | On device |
| Cross-app task automation, orders and bookings | Cloud |
| Gemini dialing businesses to complete tasks | Cloud |
| Proactive Assistance, At a Glance context | Cloud |
| Watch 5 offline Gemini | On device, limited to timers, alarms, workouts |
Everything in the local column is perception and transformation: pixels, audio, text cleanup. Real work, and the class of work the Santafe TPU was designed for. Everything in the cloud column is agency. The features Google led the keynote with, the ones that place an order or call a store on your behalf, are all round trips to a datacenter.
The Watch 5 row is the most honest line in the table. Google doubled the wearable's RAM from 2GB to 3GB specifically for Gemini, and the offline capability it earned is timers, alarms, and starting a workout. That is what a gigabyte of headroom buys at the current state of the art. Everything else on the wrist goes over the network.
Six Months, Not Twelve
The second number worth flagging is the one nobody put on a slide. Buyers of the Pro phones now get six months of bundled AI Pro access. It used to be twelve.
Read that alongside the RAM cut and the business model stops being ambiguous. If the strategy were genuinely local-first, the trial length would not matter, because the valuable features would keep working when the subscription lapsed. Google shortened the window because it expects these devices to convert into recurring cloud inference revenue, and it wants the conversion event to arrive sooner. Halving the free period on a phone whose marketing is built on AI is a forecast, not an oversight.
None of that is a scandal. Cloud inference has gotten extraordinarily cheap, the models that live there are two orders of magnitude larger than anything a handset will hold this decade, and a serving fleet can be updated on a Tuesday while a phone cannot. Routing the hard work to the datacenter is the correct engineering decision. I just want the marketing to match it. “Gemini Intelligence, embedded” describes the entry point, not the location of the model.
Where the Counterargument Sits
The strongest case against my reading is that raw capacity is the wrong metric for a phone. Distillation and quantization have moved fast enough that Nano 4 in 2026 is plausibly doing what a cloud model did in 2024, and the 3.5x energy reduction matters more to actual users than headroom for a model class that would flatten the battery in twenty minutes anyway. A phone is a thermal budget wearing a screen. Shipping a smaller, faster, cooler model that runs all day beats shipping a bigger one you throttle after two prompts. Google also spent real silicon on things that are not language models at all: a new custom accelerator for portrait video, an upgraded ISP, and a Titan M3 security chip aimed at post-quantum threats.
Fair. My response is that Google chose to frame the whole product around agentic capability, and agentic capability is where the local ceiling bites hardest. Automating a booking means holding context, calling tools, and recovering from failure across many turns. That is the workload with the worst memory profile, and it is exactly the workload Google shipped to the cloud.
The thing I find genuinely interesting is that Google is now the assistant layer on both sides of the fence. It is putting Gemini on Pixel hardware it controls and, per the WWDC arrangement, under Siri on hardware it does not. If the valuable inference happens in Google's datacenter either way, then the 2nm milestone is a margin story and a battery story, not a strategy story. The strategy already shipped, and it lives in a rack.
Three Signposts
First, whether Google publishes the memory footprint of Gemini Nano 4 or leaves it unstated. Every vendor quotes speed multiples and none of them quote resident size, because size is the number that reveals the ceiling.
Second, whether the 12GB base holds through next generation or the Pro quietly returns to 16GB once agentic features start failing on memory pressure. A reversal within a cycle would tell you this was a bill-of-materials decision that ran into physics.
Third, whether Apple takes the opposite bet in September. Apple has been the loudest voice for on-device processing and has more incentive than anyone to keep inference off a rival cloud. If the iPhone 18 ships with more memory rather than more TOPS, that is the industry conceding which number actually governs.
