본문으로 건너뛰기
All news

Gemini 3.7 Flash Takes Top Spot on Coding Benchmarks at 340 Tokens/Sec

In one line: Google launched Gemini 3.7 Flash on August 13, positioning it as the fastest and highest-scoring model on production-code benchmarks — just three weeks after its predecessor.

Key points

  • Benchmark lead: 43.6% on FrontierCode 1.1 Main, ahead of Claude Sonnet 5 (42.7%) and GPT-5.6 Terra (41.3%)
  • Coding leap: DeepSWE v1.1 jumped from 49.0% (3.6 Flash) to 65.3% — a 16-point gain in 23 days
  • Speed: ~340 output tokens/sec in independent Artificial Analysis testing, top tier among reasoning-class models
  • Pricing: $0.75 / $3.75 per million tokens (introductory; rises to $1.50 / $7.50 after December 31, 2026)
  • Context: 1M-token window unchanged from the previous version

Why it matters

Three weeks between Flash releases signals Google is on an aggressive release cadence in the coding-agent race. Pairing benchmark leadership with competitive speed and pricing makes Gemini 3.7 Flash a compelling default for agentic pipelines — though the introductory price expiry is worth factoring into long-term cost planning.

Read more

How this story unfolded

  1. xAI Launches Grok 4.6 — Better Reasoning, Same Scale, Beats Claude Opus 4.8 on Coding
  2. Anthropic Set to Make 'Auto Mode' the Default in Claude Code
  3. DeepSeek V4-Flash-0731 Official Release: Beats Its Own Pro Model on 9 Agent Benchmarks
  4. Gemini 3.5 Pro Targets July 17 Launch After Full Architecture Rebuild

Tools in this story

LLM
— Large Language Model의 약자로, '거대 언어 모델'이라고 해요. ChatGPT, Claude 같은 AI가 바로 LLM이에요. 엄청나게 많은 텍스트를 학습해서 사람처럼 글을 쓰고 대화할 수 있어요.

뉴스레터 구독

무료 뉴스레터

매주 핵심 AI 소식, 한 번에 받기

쏟아지는 AI·LLM 뉴스 중 꼭 알아야 할 것만 골라 메일로 보내드려요. 뉴스레터 발송이 시작되면 구독자분들께 가장 먼저 보내드립니다.