Gemini 3.7 Flash delivers a 78.8% relative gain on AutomationBench.
Coding performance jumps 33.3% on DeepSWE v1.1.
Lower token costs could make multi-step AI agents more economical.
Google launched Gemini 3.7 Flash on August 13, 2026, just three weeks after Gemini 3.6 Flash. The new model targets software work, AI agents, document tasks, and multi-step business jobs. Google calls it a workhorse model, with a focus on strong results at a lower cost. Reuters notes that Google still has not shared a launch date for Gemini 3.5 Pro, the expected high-end model.
Gemini 3.7 Flash targets a harder task than a normal chatbot. A chatbot can answer one request and stop. An AI agent must plan a task, use tools, read the result, fix errors, and take the next step. A long task may need many model calls.
Google reports a sharp gain on AutomationBench, where Gemini 3.7 Flash scored 30.4%, up from 17.0% for Gemini 3.6 Flash. That marks a 78.8% relative rise. Google says the model can handle roadblocks, follow detailed instructions, plan tasks, and use tools with better control.
Google reports a 43.6% score for Gemini 3.7 Flash on FrontierCode 1.1 Main, versus 34.4% for Gemini 3.6 Flash. That marks a 26.7% relative rise.
DeepSWE v1.1 shows an even larger jump. Gemini 3.7 Flash scored 65.3%, compared with 49.0% for Gemini 3.6 Flash. That equals a 33.3% relative rise. DeepSWE tests complex software tasks, so the result gives a useful view of code work.
GitHub has also added Gemini 3.7 Flash to Copilot. Early GitHub tests show gains in web and app development, agent code work, code quality, codebase research, result checks, and final output quality.
Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. From January 1, 2027, the list price rose to $1.50 per million input tokens and $7.50 per million output tokens.
An agent may call a model many times for one task. Lower token cost can reduce the bill, while better results can reduce repeat calls. That makes price a major part of Google’s strategy, not just a side benefit.
Also Read - Best AI Browser Assistants for Daily Work in 2026
A good agent must understand a goal, choose a tool, inspect a result, spot an error, and move ahead. Fewer failed steps can cut total cost even if raw speed stays close to the old model.
Google has placed Gemini 3.7 Flash across its own agent stack. The model supports the Gemini API, Google AI Studio, Google Antigravity, Android Studio, and Gemini Enterprise Agent Platform. Gemini Spark also uses it across more than 160 countries.
Gemini 3.7 Flash scored 1,588 Elo on WebDev Arena, up from 1,538 for Gemini 3.6 Flash. That 50-point rise equals about 3.3%. The smaller gain still gives a useful web test.
Google says the model can create more complete web apps with fewer prompts and can follow screenshots, layouts, and design systems with better accuracy.
The main benchmark figures come from Google’s own tests. They show clear gains, but independent tests still matter. Agent quality also depends on tools, context, memory, system design, and the cost of failed actions.
The next few weeks should give a clearer view. Developers will test the model across real code bases and long agent tasks. Those results will show whether the lower price and better scores translate into lower cost per successful task.
Also Read - Top 10 AI Chatbot Development Companies in India (2026)
Gemini 3.7 Flash shows a clear strategy. Google wants capable AI that can run many steps at a low price. The aim goes beyond a smarter chatbot. The bigger goal centers on AI agents that can handle software work, business tasks, and digital chores with less human help.
The key figures tell that story: a 50% lower launch price, a 78.8% relative gain on AutomationBench, a 33.3% rise on DeepSWE v1.1, and a 26.7% rise on FrontierCode 1.1 Main. If real-world tests support those results, Gemini 3.7 Flash could give Google a strong tool for the next phase of AI: cheap, fast, capable agents that can finish useful work.
What is Gemini 3.7 Flash?
Google’s latest AI model focused on agents, coding, documents, and multi-step workflows.
Why is it important for AI agents?
It is designed to plan, use tools, handle roadblocks, and complete longer tasks more effectively.
How much does Gemini 3.7 Flash cost?
It launches at $0.75 per million input tokens and $3.75 per million output tokens through December 2026.
Is Gemini 3.7 Flash better for coding?
Google reports significant gains across software engineering benchmarks, including DeepSWE and FrontierCode.
What could determine its real-world success?
Independent testing will show whether benchmark improvements translate into lower cost and better performance on real agent tasks.