Google, OpenAI, and DeepSeek all shipped in about 24 hours, and none of them competed on “smarter.” One cut the price in half and made it available everywhere. One put its flagship on a dinner-plate chip and hit 750 tokens a second. One left preview, added a thinking dial, and warned that prices jump Sunday.
The race this week is latency, cost, and whether you can actually use the thing. Here is who to pick this weekend.
Gemini 3.7 Flash: the one you can use today
Google shipped Gemini 3.7 Flash on August 13, three weeks after 3.6 Flash. It is the workhorse for coding and agents, not the missing flagship. Gemini 3.5 Pro still does not have a date.
On Google’s own table vs 3.6 Flash: FrontierCode 1.1 Main 43.6% vs 34.4%, DeepSWE v1.1 65.3% vs 49.0%, WebDev Arena Elo 1588 vs 1538, GDP.pdf 34% vs 22%, AutomationBench 30.4% vs 17%. The useful claim is fewer retries and better first-pass code, not a new intelligence class.
The price is the product. Introductory rate through December 31: $0.75 / $3.75 per million input/output tokens, half of 3.6 Flash. On January 1 it goes to $1.50 / $7.50. It is live in the Gemini API, AI Studio, Android Studio, Google Antigravity, Gemini Enterprise, and inside Spark for AI Pro and Ultra subscribers in 160+ countries.
Use it if: you want a cheaper coding/agent model this afternoon and you do not want a waitlist.
GPT-5.6 Sol Ultrafast: same brain, rented speed
OpenAI did not ship a new model. Ultrafast is GPT-5.6 Sol served on Cerebras wafer-scale chips: up to 750 output tokens per second, about 14x standard processing, no quality cut. Cerebras, citing Artificial Analysis, puts that at 11x Claude Fable 5 and 5x Opus 4.8 Fast mode.
Cerebras also ran Humanity’s Last Exam: Sol Ultrafast finished all 2,500 questions in 11 hours 11 minutes. Fable 5 took 78 hours 27 minutes at comparable accuracy. On GDP-Val they claim a 5.6x end-to-end speedup with no quality drop. Those are Cerebras’s benches, not a third-party bake-off. Treat them as a speed claim, not a ranking.
The hardware story matters. Cerebras puts 44 GB of SRAM on one wafer-sized chip so weights stay on-die instead of shuttling off a GPU. OpenAI still trains on Nvidia. For this tier, it is renting inference from someone else. No public price. Limited preview, waitlist, Jane Street and Podium already on it for voice and markets work.
Use it if: latency is the product (voice, incident response, live support) and you can get on the list. Do not wait for it if you need something in production tonight.
DeepSeek V4-Pro: cheap, huge context, prices move Sunday
DeepSeek flipped deepseek-v4-pro to the 0813 general-availability build. Preview started in April. No splashy blog post. The pricing page is the announcement.
What you get now: 1 million tokens of context, 384K max output, thinking or non-thinking, plus low / high / max effort. OpenAI-format, Anthropic-format, and a Responses API with Codex support. DeepSeek’s own agent table: Terminal Bench 2.1 87.9, DeepSWE 62.7, NL2Repo 61.5. Independent writeups put it a few points behind Claude Fable 5 on a basket of agent benches, at a tiny fraction of the price.
Current API: $0.435 / $0.87 per million input/output, cache hits at $0.003625. Concurrency cap is 500 (Flash is 2,500).
That cheap window closes. At 16:00 UTC on August 16 (noon ET), DeepSeek switches to peak / off-peak. Peak hours are 01:00–04:00 and 06:00–10:00 UTC. V4-Pro becomes $0.66 / $1.98 off-peak and $1.32 / $3.96 at peak. Still cheaper than Flash’s intro rate, let alone Fable 5 at $10 / $50. If you have a fat agent loop to run, do it before Sunday.
Use it if: you want a million-token context and you can dial thinking to the task. Check the clock. After Sunday, time-of-day is part of the bill.
Who to use this weekend
- Need it now, coding or agents, no waitlist: Gemini 3.7 Flash. Half price through December.
- Need frontier quality at voice speed, and you can get invited: Sol Ultrafast. Same model, rented silicon, no public price.
- Need a huge context window and a thinking dial, and you can run the job before Sunday: DeepSeek V4-Pro at current rates. After 16:00 UTC August 16, budget for peak hours.
None of this replaces Claude or Grok if those already fit. It is the other axis. Yesterday’s story was long-running agents. Today’s is whether those agents answer before you context-switch. Google made that cheap and available. OpenAI made it fast and scarce. DeepSeek made it wide, then put a timer on the discount.
Related: Yesterday’s Grok 4.6 / DeepMind / Hugging Face roundup.
Share this article

Leave a Reply