본문으로 건너뛰기
All news

xAI Releases Grok Voice Think Fast 2.0: 0.70s Latency, 60% Fewer Reasoning Tokens

Summary: xAI's Grok Voice Think Fast 2.0 cuts time-to-first-audio to 0.70 seconds — a 44% improvement over its predecessor — while introducing a parallel reasoning architecture that runs inference concurrently with speech output.

Key Facts

  • 0.70s first-audio latency: Down from 1.25s for the prior Think Fast model, matching the natural pause length in human conversation.
  • 60% fewer inference tokens: Greater token efficiency means tool calls typically complete before the agent finishes its first spoken sentence.
  • Reason while speaking: The model processes the next reasoning step in parallel with its audio output rather than sequentially, improving effective intelligence without adding latency.
  • Default swap: From August 5, the grok-voice-latest API alias automatically routes to Think Fast 2.0.

Why It Matters

Perceived latency is the primary adoption barrier for voice AI agents. Research suggests user drop-off falls sharply once response time dips below one second — the threshold that makes AI feel conversational rather than robotic. Grok Voice 2.0 now sits comfortably inside that range. The concurrent-reasoning design is a notable architectural bet: if it holds up under independent benchmarking, it could give xAI a durable edge in agentic voice applications, where both speed and reasoning depth matter.

More

뉴스레터 구독

무료 뉴스레터

매주 핵심 AI 소식, 한 번에 받기

쏟아지는 AI·LLM 뉴스 중 꼭 알아야 할 것만 골라 메일로 보내드려요. 뉴스레터 발송이 시작되면 구독자분들께 가장 먼저 보내드립니다.