Google shipped Gemini 3.6 Flash on July 21 — not the flagship everyone was waiting on, which is still in partner testing. Output dropped from $9 to $7.50 per million tokens, and the model uses about 17% fewer output tokens to do the same work. Those stack: cheaper tokens, fewer of them. On the composite intelligence index it's roughly flat. This wasn't a smarter model. It was a cheaper one.
Worth reading the numbers carefully. The 17% is measured across the whole index; the 65% figure making the rounds comes from specific benchmarks under specific conditions. Budget on the first number, not the second. The quieter upgrade may matter more day to day: the model's knowledge cutoff finally moved from January 2025 to March 2026.
Here's why a price cut is the story rather than a footnote. An agent that runs on a schedule — pulling a weekly report, drafting a follow-up sequence, monitoring a listing feed — has a monthly bill, and that bill is what decides whether it stays switched on. Every vendor is now competing on cost per completed task rather than benchmark scores, which is what happens when a technology stops being a demo and starts being infrastructure. For anyone out here running something real on top of these models, that competition is the part worth tracking. California's half-price deal for cities and counties was the same story wearing a different hat.
The frontier gets the headlines. The price per task decides what actually gets built.