Sunday’s hike is not leftover weekend news. It is already on today’s invoice. DeepSeek flipped the API to peak and off-peak rates at 16:00 UTC on Sunday, August 16 (12:00 PM ET). Monday, August 17 is the first full US business day on the new card.
Peak is not your afternoon standup. In Eastern Time it lands overnight. If a cron job still treats 2 a.m. as the cheap window, that assumption just got expensive.
Source: DeepSeek’s August 13 GA note and the live Models & Pricing page.
The ET schedule
Official peak hours are 01:00–04:00 and 06:00–10:00 UTC. Every other hour is off-peak. Off-peak is half of peak. DeepSeek writes the same windows in Beijing time (9:00–12:00 and 14:00–18:00).
In August (EDT, UTC-4) peak is:
- 9:00 PM–12:00 AM ET
- 2:00 AM–6:00 AM ET
US weekday work hours sit entirely in off-peak. A 10 a.m. New York call is the cheap rate. A 3 a.m. batch job is not.
Do not invert this. DeepSeek’s peak tracks Beijing work hours, not yours.
The official table
US dollars per million tokens, copied from DeepSeek’s official Models & Pricing card on August 17, 2026. Off-peak is exactly half of peak on every line.
| Model | Window | Cache hit | Cache miss | Output |
|---|---|---|---|---|
| V4-Flash | Off-peak | $0.007 | $0.22 | $0.66 |
| V4-Flash | Peak | $0.014 | $0.44 | $1.32 |
| V4-Pro | Off-peak | $0.022 | $0.66 | $1.98 |
| V4-Pro | Peak | $0.044 | $1.32 | $3.96 |
The old flat rates are no longer on the live page. Before 16:00 UTC on August 16, V4-Pro billed about $0.435 miss / $0.003625 hit / $0.87 out at every hour (the card we cited in Friday’s post). V4-Flash billed about $0.14 / $0.0028 / $0.28. Treat both rows as pre-August 16 flats, not as anything you will find on the live page today.
Even off-peak is a raise against that old card. Peak is 2× off-peak, not 2× the old rate. V4-Pro output at peak ($3.96) is about 4.6× the old $0.87 flat. Off-peak ($1.98) is about 2.3× the old flat.
The cache trap
The line that will surprise agent workloads is cache-hit input.
On V4-Pro, a cache hit used to cost $0.003625 per million tokens. It is now $0.022 off-peak and $0.044 at peak. Using only those sourced numbers, that is about 6× off-peak and about 12× at peak.
The relative discount survived. A hit is still about 3% of a miss ($0.022 vs $0.66 off-peak). The absolute price is not “nearly free” anymore. If you resend a large system prompt on every agent step, that prefix is now a real line item.
Flash moved too, from a $0.0028 flat hit to $0.007 off-peak and $0.014 peak. Smaller multiple. Still not free.
What to change this week
Audit cron and overnight agents first. The expensive windows in ET are 9:00 PM–midnight and 2:00–6:00 AM. Daytime US is cheaper. Do not run overnight batch or agent jobs in those two windows unless you mean to pay peak.
Keep Flash as the cheap lane. Same clock, one-third the output rate: $0.66 off-peak and $1.32 at peak, versus Pro at $1.98 and $3.96. Use Pro when the task needs it. Do not default every loop to Pro because last week’s bill looked tiny.
Stop assuming cache is nearly free. Check prompt_cache_hit_tokens on a few real calls this week. If a fat prefix is hitting, recost it at $0.022 (or $0.044) per million, not at $0.003625.
One cautious comparison, from Google’s August 13 Gemini 3.7 Flash intro card ($0.75 input / $3.75 output): DeepSeek Flash off-peak output at $0.66 is still lower. That is not a claim that DeepSeek undercuts every Western model at every hour. Re-check the other labs’ live cards before you say “still cheapest.”
Monday checklist
- Peak in ET is 9:00 PM–12:00 AM and 2:00–6:00 AM. Daytime US is off-peak.
- Move overnight batch and agent jobs out of those two windows.
- Keep Flash for high-volume loops. Use Pro on purpose.
- Recost anything that leaned on cache hits at $0.003625.
- Re-read the live pricing page before you lock a budget. DeepSeek says the card can move.
- If a job can wait until 6:00 AM ET, wait.
Share this article

Leave a Reply