OpenAI Deploys GPT-5.6 Sol on Cerebras at Up to 750 Tokens Per Second
Summary: OpenAI confirmed it will run GPT-5.6 Sol on Cerebras wafer-scale hardware in July at up to 750 tokens per second — roughly an order of magnitude faster than typical GPU-based frontier inference.
Key Facts
- Cerebras throughput: up to 750 t/s for Sol; standard GPU clusters serve frontier models at 40–120 t/s
- Pricing: Sol at $5 input / $30 output; Terra at $2.50/$15; Luna at $1/$6 (per 1M tokens)
- Rollout limited to ~20 select partners initially, following the White House's voluntary pre-deployment safety review request
- Sol targets agentic, long-horizon, and scientific reasoning tasks; Terra and Luna are optimized for cost-sensitive general workloads
Why It Matters
Serving a frontier model at 750 t/s collapses multi-step agent loops from overnight jobs to real-time interactions. For developers building coding assistants, research agents, and complex orchestration pipelines, this speed tier removes the biggest practical bottleneck — waiting for the model — and resets expectations for what interactive AI tooling can do.
More
How this story unfolded
- GPT-5.6 Sol Broad Release Targeting Mid-July After Government-Gated Preview
- OpenAI's GPT-5.6 Launches in Three Tiers: Sol's 'Ultra Mode' Bets on Multi-Agent Orchestration
- OpenAI Launches GPT-5.6 as Sol, Terra, and Luna — Restricted to US Trusted Partners
- GPT-5.6 June Window Closes — Prediction Market Odds Collapse From 83% to 18%
- LLM
- — Large Language Model의 약자로, '거대 언어 모델'이라고 해요. ChatGPT, Claude 같은 AI가 바로 LLM이에요. 엄청나게 많은 텍스트를 학습해서 사람처럼 글을 쓰고 대화할 수 있어요.