OpenAI and Broadcom Unveil Jalapeño, a Custom LLM Inference Chip
One-liner: OpenAI and Broadcom introduced Jalapeño, OpenAI's first custom AI accelerator — a reticle-sized ASIC designed from the ground up for LLM inference, not repurposed from training hardware.
Key Facts
- Full-reticle ASIC purpose-built for LLM inference; designed in nine months using OpenAI's own AI models to accelerate chip development
- Developed with Broadcom (manufacturer) and Celestica (production scale-up) — early testing shows better performance per watt than current state-of-the-art chips
- Microsoft will purchase 40% of initial production; full deployment targeted for late 2026
- OpenAI aims for 10 gigawatts of Jalapeño-powered compute by 2029, roughly the output of ten nuclear reactors
Why It Matters
This is OpenAI's clearest move yet toward hardware self-sufficiency at the inference layer. Training still depends on Nvidia GPUs, but inference is where serving costs compound at scale — every ChatGPT query runs it. A purpose-built inference ASIC could meaningfully reshape OpenAI's unit economics as it approaches its IPO, while signaling that Nvidia's dominance at the edge of the AI stack is no longer guaranteed.
Read More
How this story unfolded
- OpenAI Launches GPT-5.5-Cyber and Patch the Planet to Automate Open-Source Vulnerability Fixes
- Open-Source AI Startup Reflection Signs $6.3B Compute Deal With SpaceX
- Amazon Eyes Direct Trainium Sales to Rival Nvidia in Data Center Market
- NVIDIA Kicks Off Blackwell Ultra B300 Mass Production — 50x Hopper Throughput per Watt
- LLM
- — Large Language Model의 약자로, '거대 언어 모델'이라고 해요. ChatGPT, Claude 같은 AI가 바로 LLM이에요. 엄청나게 많은 텍스트를 학습해서 사람처럼 글을 쓰고 대화할 수 있어요.