Gemini 3.7 Flash Takes Top Spot on Coding Benchmarks at 340 Tokens/Sec
In one line: Google launched Gemini 3.7 Flash on August 13, positioning it as the fastest and highest-scoring model on production-code benchmarks — just three weeks after its predecessor.
Key points
- Benchmark lead: 43.6% on FrontierCode 1.1 Main, ahead of Claude Sonnet 5 (42.7%) and GPT-5.6 Terra (41.3%)
- Coding leap: DeepSWE v1.1 jumped from 49.0% (3.6 Flash) to 65.3% — a 16-point gain in 23 days
- Speed: ~340 output tokens/sec in independent Artificial Analysis testing, top tier among reasoning-class models
- Pricing: $0.75 / $3.75 per million tokens (introductory; rises to $1.50 / $7.50 after December 31, 2026)
- Context: 1M-token window unchanged from the previous version
Why it matters
Three weeks between Flash releases signals Google is on an aggressive release cadence in the coding-agent race. Pairing benchmark leadership with competitive speed and pricing makes Gemini 3.7 Flash a compelling default for agentic pipelines — though the introductory price expiry is worth factoring into long-term cost planning.
Read more
- Gemini 3.7 Flash Model Card — Google DeepMind
- Gemini 3.7 Flash: Pricing, Benchmarks and What Changed — FelloAI
How this story unfolded
- xAI Launches Grok 4.6 — Better Reasoning, Same Scale, Beats Claude Opus 4.8 on Coding
- Anthropic Set to Make 'Auto Mode' the Default in Claude Code
- DeepSeek V4-Flash-0731 Official Release: Beats Its Own Pro Model on 9 Agent Benchmarks
- Gemini 3.5 Pro Targets July 17 Launch After Full Architecture Rebuild
Tools in this story
- LLM
- — Large Language Model의 약자로, '거대 언어 모델'이라고 해요. ChatGPT, Claude 같은 AI가 바로 LLM이에요. 엄청나게 많은 텍스트를 학습해서 사람처럼 글을 쓰고 대화할 수 있어요.