본문으로 건너뛰기
All news

OpenAI Deploys GPT-5.6 Sol on Cerebras at Up to 750 Tokens Per Second

Summary: OpenAI confirmed it will run GPT-5.6 Sol on Cerebras wafer-scale hardware in July at up to 750 tokens per second — roughly an order of magnitude faster than typical GPU-based frontier inference.

Key Facts

  • Cerebras throughput: up to 750 t/s for Sol; standard GPU clusters serve frontier models at 40–120 t/s
  • Pricing: Sol at $5 input / $30 output; Terra at $2.50/$15; Luna at $1/$6 (per 1M tokens)
  • Rollout limited to ~20 select partners initially, following the White House's voluntary pre-deployment safety review request
  • Sol targets agentic, long-horizon, and scientific reasoning tasks; Terra and Luna are optimized for cost-sensitive general workloads

Why It Matters

Serving a frontier model at 750 t/s collapses multi-step agent loops from overnight jobs to real-time interactions. For developers building coding assistants, research agents, and complex orchestration pipelines, this speed tier removes the biggest practical bottleneck — waiting for the model — and resets expectations for what interactive AI tooling can do.

More

How this story unfolded

  1. GPT-5.6 Sol Broad Release Targeting Mid-July After Government-Gated Preview
  2. OpenAI's GPT-5.6 Launches in Three Tiers: Sol's 'Ultra Mode' Bets on Multi-Agent Orchestration
  3. OpenAI Launches GPT-5.6 as Sol, Terra, and Luna — Restricted to US Trusted Partners
  4. GPT-5.6 June Window Closes — Prediction Market Odds Collapse From 83% to 18%
LLM
— Large Language Model의 약자로, '거대 언어 모델'이라고 해요. ChatGPT, Claude 같은 AI가 바로 LLM이에요. 엄청나게 많은 텍스트를 학습해서 사람처럼 글을 쓰고 대화할 수 있어요.

뉴스레터 구독

무료 뉴스레터

매주 핵심 AI 소식, 한 번에 받기

쏟아지는 AI·LLM 뉴스 중 꼭 알아야 할 것만 골라 메일로 보내드려요. 뉴스레터 발송이 시작되면 구독자분들께 가장 먼저 보내드립니다.