Gemini 3.7 Flash Is Cheap Until New Year. Plan for January Now

Google's new workhorse model is a real step up for coding agents, but its introductory price doubles on 1 January.

2 min read ·

Google released Gemini 3.7 Flash on 13 August, generally available on day one through the Gemini API, AI Studio, Android Studio and Google Antigravity. Google pitches it as its best workhorse model for coding and agents, and it arrives only three weeks after Gemini 3.6 Flash. It keeps a context window of 1,048,576 tokens with up to 65,536 tokens of output.

I build products on these models for a living, so here is what I think matters if you ship software, rather than what topped which chart.

The jump is real, if Google's numbers hold

Google's reported gains over 3.6 Flash are not incremental. On DeepSWE v1.1, a benchmark of software engineering tasks in real codebases, the model scores 65.3 per cent against 49.0 per cent for its predecessor. On FrontierCode it moves from 34.4 to 43.6 per cent, and on AutomationBench, which tests multi-step business workflows, from 17.0 to 30.4 per cent, according to DataNorth's summary of the launch.

These are vendor numbers. Still, an AutomationBench score that nearly doubles is the kind of change you can feel in an agent: fewer stalled tool loops, fewer half-finished tasks. If you have a Flash-class model handling routing, extraction or first-pass code changes, it is worth running your own evals this week.

Read the price carefully

The launch price is $0.75 per million input tokens and $3.75 per million output tokens. That is introductory. From 1 January 2027 it becomes $1.50 and $7.50, exactly double. Context caching costs $0.075 per million tokens plus $0.50 per million tokens per hour of storage, and both of those double in January too.

Google's own August round-up frames the price as half the original 3.6 Flash cost, which is true today. The risk is that teams build a margin model in August on the August price and get a nasty surprise in their January invoice. A few practical habits:

  • Cost your features at the January rate. If a feature only makes sense at $0.75 input, it does not make sense.
  • Use caching aggressively. Agents resend the same system prompt and tool definitions constantly. The storage fee is small next to repeated input tokens.
  • Put a router in front of the model. If you call models through a thin abstraction, switching in January is a config change, not a sprint.

Three-week release cycles change how you test

The bigger shift for builders is cadence. A new Flash every three weeks means the model under your product is no longer a fixed dependency. Each upgrade can improve one prompt and quietly break another. The answer is boring and necessary: keep a small eval set of your real tasks, with expected outputs, and run it every time a model changes. Twenty representative cases you trust beat a public leaderboard you do not.

One regional note. Consumer access to 3.7 Flash comes through Spark, Google's personal agent, for AI Pro and Ultra subscribers, but not in the European Economic Area, the UK, Switzerland or Nigeria. Malaysia and the rest of Southeast Asia are not on that exclusion list, so teams here with AI Pro or Ultra can try it in Spark as well as through the API.

My take: Gemini 3.7 Flash is a strong default for high-volume agent work right now. Just treat the next four and a half months as a discount period, not the price.


Sources

Responses (1)

Sign in to leave a response.

  • The eval-set point deserves its own post. We pinned a model version last year purely because nobody could tell us whether the new one was better for our prompts.

More from Fakhrul

Recommended from Horizon