본문으로 건너뛰기
All news

'Neuralese': The Worry That AI May Think in a Language We Can't Read

In one line: "Neuralese" describes AI reasoning that happens through internal numeric representations rather than human-readable text — and the growing concern is that it makes a model's thought process much harder to monitor.

Key points

  • Neuralese reportedly refers to models passing their chain-of-thought as high-dimensional vectors or latent representations instead of human-language tokens.
  • Today's reasoning models mostly leave a step-by-step trace in readable text; shifting to reasoning directly in latent space would remove that legible trail.
  • Safety researchers warn that if reasoning moves to neuralese, oversight techniques like chain-of-thought monitoring could be rendered ineffective.

Why it matters

Much of AI safety rests on the assumption that a human can read and inspect why a model produced an answer. If reasoning runs through an internal language people cannot decode, spotting dangerous intent or malfunction in advance becomes far harder — which is the crux of the concern.

Read more

뉴스레터 구독

무료 뉴스레터

매주 핵심 AI 소식, 한 번에 받기

쏟아지는 AI·LLM 뉴스 중 꼭 알아야 할 것만 골라 메일로 보내드려요. 뉴스레터 발송이 시작되면 구독자분들께 가장 먼저 보내드립니다.