In partnership with

Nine months. That's how long it took OpenAI and Broadcom to build Jalapeño — OpenAI's first custom chip for running AI models. A typical custom chip takes two to five years from design to tape-out. OpenAI did it in nine months, partly by using AI to help design it. Broadcom called it "what may be the fastest ASIC development cycle ever."

Jalapeño is an inference chip — designed for running AI models, not training them. Training is the expensive one-time process that builds a model from scratch (months of compute, hundreds of millions of dollars). Inference is what happens every time you type a message and get a response back — constant, high-volume, cost-per-query. Every ChatGPT response runs through inference hardware. OpenAI has been running all of that on Nvidia. Jalapeño changes that for the inference side. Pre-training still runs on Nvidia GPUs.

The nine-month timeline is partly possible because AI models helped accelerate the chip design process itself. That feedback loop — AI helping build the infrastructure that runs AI — is the thing worth tracking. The chip was built for gigawatt-scale deployment with Microsoft in 2026. On performance, OpenAI says Jalapeño delivers performance per watt substantially better than current state-of-the-art inference hardware. Detailed benchmarks haven't been published yet.

Make Tax Season Simple

Tax season doesn't have to mean wondering if you have the right forms, second-guessing your deductions, or scrambling to pull everything together before the deadline.

With BELAY’s tax prep support, you can approach tax season with confidence. Stay organized with one centralized place to gather and check off your documents, keep track of valuable deductions like HSA contributions and education expenses while leaning on experienced professionals who make tax preparation accurate, efficient, and completely hands-off.

Download BELAY's free Personal Tax Checklist and start preparing with confidence, today.

OpenAI has been almost entirely dependent on Nvidia for compute. That's a real constraint — it affects pricing, capacity, and roadmap. A custom inference chip means OpenAI controls the full stack for the product they actually sell. Whether this efficiency gain reaches you as lower API prices is a business decision, not a guaranteed outcome. OpenAI hasn't committed to specific price cuts tied to Jalapeño. Secondary coverage that circulated specific percentage savings versus Nvidia didn't appear in the primary sources — that claim was dropped.

The honest version: inference is the main ongoing cost of running a large language model. Jalapeño appears to reduce that cost significantly per watt. What OpenAI does with those savings is still up to OpenAI. What's already changed is that they no longer need Nvidia's permission to control their own inference stack.

If you want to sound smarter without doing more…

Read Morning Brew. 5-minute daily email. Day's most important news. Explained in simplest terms. Really, it’s a news experience you’ll actually enjoy.

At this point, it's not about trying Morning Brew's daily newsletter—It's about missing out.

Click below to join 4,000,000+ daily readers.

— AI Super Simplified, Edition 292

Keep Reading