본문으로 건너뛰기
All news

Anthropic Proposes Five-Tier Industry Standard for Rating AI Jailbreak Severity

Summary: Anthropic published the Cyber Jailbreak Severity (CJS) framework — a five-tier, exponential scale co-developed with Amazon, Microsoft, and Google to standardize how dangerous an AI jailbreak is rated.

Key Points

  • Five tiers from CJS-0 (Informational) to CJS-4 (Critical); each step represents exponentially greater risk than the last
  • Severity is scored across four axes: capability gain, breadth of harm, ease of weaponization, and discoverability
  • Scope is limited to cybersecurity jailbreaks; non-cyber misuse (e.g., system-prompt extraction) is explicitly excluded
  • Released alongside a new Claude Fable 5 safety classifier that Anthropic says blocks over 99% of attempts replicating the June jailbreak that triggered a 19-day US export-control suspension

Why It Matters

Until now, there has been no shared vocabulary for comparing jailbreak incidents across the AI industry. If adopted, CJS gives companies, regulators, and security researchers a common scale — analogous to CVSS scores in traditional cybersecurity — to communicate incident severity and set consistent remediation thresholds.

Read More

How this story unfolded

  1. OpenAI Launches GPT-5.6 as Sol, Terra, and Luna — Restricted to US Trusted Partners
  2. OpenAI Launches GPT-5.5-Cyber and Patch the Planet to Automate Open-Source Vulnerability Fixes
  3. Five Eyes Agencies: AI-Powered Cyberattacks Are Months, Not Years, Away
  4. Claude Mythos Uncovers 10,000+ Critical Software Flaws in First Month
LLM
— Large Language Model의 약자로, '거대 언어 모델'이라고 해요. ChatGPT, Claude 같은 AI가 바로 LLM이에요. 엄청나게 많은 텍스트를 학습해서 사람처럼 글을 쓰고 대화할 수 있어요.

뉴스레터 구독

무료 뉴스레터

매주 핵심 AI 소식, 한 번에 받기

쏟아지는 AI·LLM 뉴스 중 꼭 알아야 할 것만 골라 메일로 보내드려요. 뉴스레터 발송이 시작되면 구독자분들께 가장 먼저 보내드립니다.