Migration scope
Google released Gemini 3.7 Flash as a generally available model for coding, agentic workflows, and multimodal reasoning. Its stable API model ID is gemini-3.7-flash.
Google lists Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3 Flash Preview, and Gemini 3.1 Pro as migration sources. This guide summarizes Google's published migration requirements. ModelRun Lab has not independently tested application compatibility.
1. Change the model ID
Update the model string in your application to:
gemini-3.7-flash
Do not rely on an unversioned alias when a production application needs a predictable migration target. Confirm regional and account availability in the official model documentation before rollout.
2. Remove deprecated sampling parameters
Google instructs developers moving from older Gemini models to remove these generation parameters:
temperaturetop_ptop_kcandidate_count
Gemini 3.6 Flash already did not support the first three parameters. Applications moving directly from 3.6 may therefore have completed part of this cleanup already.
3. Replace thinking budget with thinking level
Replace numeric thinking_budget configuration with the string-based thinking_level option. Gemini 3.7 Flash supports low, medium, and high thinking levels, with medium as the documented default.
The appropriate level is workload-specific. Lower effort can reduce latency, while higher effort may spend more time on complex reasoning. Test the choice against your own latency, quality, and cost requirements.
4. Remove prefilled model turns
Google requires developers to remove prefilled model turns. Multi-turn interactions should use the server-side previous_interaction_id mechanism documented for the Interactions API.
This can affect applications that previously constructed conversation history by inserting assistant content manually. Review stored conversation state and replay logic before switching production traffic.
5. Audit function calling
Google's checklist calls out several function-calling changes:
- Place multimodal assets inside the response payload.
- Format inline instructions with clear double-newline separation.
- For the generateContent API, ensure each FunctionResponse includes both
call_idandname. - Investigate pre-tool text when Malformed_Function_Call errors appear.
These requirements are especially important for agents that mix tool calls, images, files, or structured outputs.
6. Check context and output limits
Google documents a one-million-token context window and a 64,000-token maximum output for Gemini 3.7 Flash. Large limits do not remove the need for application-side controls: cap untrusted input, validate tool arguments, and set practical output limits for cost and latency.
7. Run a staged migration
A safe rollout sequence is:
- Update the SDK and model ID in a development environment.
- Remove unsupported parameters and prefilled turns.
- Replay a representative evaluation set.
- Test tool calls and multimodal payloads separately.
- Compare latency, errors, output quality, and cost.
- Send a small percentage of production traffic to the new model.
- Expand only after error rates and task quality remain acceptable.
This staged rollout is editorial operational guidance, not a Google guarantee.
Evidence boundaries
- Model ID, supported thinking levels, context limits, and migration requirements are taken from Google documentation.
- Rollout advice is ModelRun Lab editorial guidance.
- ModelRun Lab has not independently benchmarked Gemini 3.7 Flash.
- Pricing and API behavior may change; confirm the official page before deployment.