본문으로 건너뛰기
All news

OpenAI Previews 'Ultrafast' Mode for GPT-5.6 Sol — Up to 14x Faster via Cerebras

In one line: OpenAI and Cerebras are previewing an "Ultrafast" API tier that runs GPT-5.6 Sol at up to 750 output tokens per second — 14 times faster than the standard tier.

Key points

  • The Ultrafast tier targets 750 output tokens per second, powered by Cerebras silicon, which OpenAI positions as up to 14× its Standard tier.
  • It is in limited API preview for a select group of customers; no price, no GA date, and no model ID string have been published yet.
  • Initial use cases being tested include coding assistance, finance, customer support, research, and incident response.
  • All performance figures are OpenAI-measured; no independent third-party benchmark exists yet.

Why it matters

Latency is one of the hardest constraints for real-time AI agents and interactive services. If the 14× speed claim holds up under independent testing, it opens the door for AI adoption in domains that have been locked out by lag — live customer interactions, code autocomplete, and real-time analytics. The missing price tag, however, means total cost of ownership is still unknown.

Read more

How this story unfolded

  1. OpenAI's GPT-5.6 Sol Hits Up to 14× Faster Responses in 'Ultrafast' Mode
  2. OpenAI Cuts GPT-5.6 Luna Price 80% — Chinese AI Competition Forces the Move
  3. DeepSeek API Breaking Change: Migrate to deepseek-v4-pro Before July 24
  4. OpenAI Deploys GPT-5.6 Sol on Cerebras at Up to 750 Tokens Per Second
API
— Application Programming Interface의 약자예요. 개발자가 AI 기능을 자기 앱에 연결할 때 사용하는 방법이에요. 일반 사용자는 몰라도 되지만, 앱 개발할 때 필요해요.

뉴스레터 구독

무료 뉴스레터

매주 핵심 AI 소식, 한 번에 받기

쏟아지는 AI·LLM 뉴스 중 꼭 알아야 할 것만 골라 메일로 보내드려요. 뉴스레터 발송이 시작되면 구독자분들께 가장 먼저 보내드립니다.