OpenAI made a chip 'and it is fast.' Volume waits for 2027
2026-08-27 · 4 min read
On Tuesday OpenAI published the first benchmark numbers for Jalapeño, the inference chip it designed with Broadcom and first announced in June. @sama summed up the launch in eight words: "we made a chip and it is fast."
The official @OpenAI thread took a few more. It says the chip delivers "more intelligence from every watt and faster responses," with higher throughput and lower latency in one architecture. A second post in the thread says OpenAI plans to begin deploying Jalapeño in its own compute infrastructure by year-end, with a Gen 2 already deep in development and a Gen 3 taking shape.
Then came the analyst who had been in the room. @dylan522p, who runs SemiAnalysis, wrote that "usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin," and said his team was let into OpenAI's lab to take the chip apart.
The numbers
OpenAI ran Jalapeño through SemiAnalysis's InferenceX benchmark on three open-weight models: GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T. Against Nvidia's GB200 and GB300 rack systems, it reports 1.5 to 1.9 times more work per watt at peak throughput, 1.7 to 3.6 times lower end-to-end latency, and 2.1 to 4.1 times higher performance on the most interactive workloads, the kind where a person is waiting on every token.
Richard Ho, who leads hardware at OpenAI, described the design goal as minimizing data movement and communication delays. The pitch is that you no longer have to pick between serving many users cheaply and serving one user quickly.
What the numbers leave out
SemiAnalysis's full write-up is more careful than the headline post. Four things stand out:
- Every number came from OpenAI. SemiAnalysis checked them during its visit to OpenAI's lab, not on hardware it controls.
- Blackwell is the wrong yardstick. Jalapeño uses HBM4 memory, so the fair fight is Nvidia's Vera Rubin. SemiAnalysis calls the Blackwell comparison "somewhat incomplete and unfair" and says that against Rubin the two land at almost the same output tokens per dollar.
- The tests were short and single-turn. SemiAnalysis's AgentX benchmark, which mimics long, multi-turn agent workloads, was not run. Jalapeño's runs also skipped speculative decoding, a standard production trick, so the gap could move in either direction once it is switched on.
- Only engineering samples exist. Production ramps through 2027, with most of the output scheduled for the end of next year.
None of that makes the chip a dud. Holding Rubin to a draw on cost per token with a first-generation design is a real result, and OpenAI says it will keep deploying Nvidia and other accelerators anyway. But the eight-word version and the long version tell different stories, and only one of them is about something that exists in volume.
My read
The interesting number is the calendar, not the speedup. Design work started in mid-2024, the chip taped out in November 2025, the first numbers arrived this week, and volume comes at the end of 2027. That is more than three years for one generation, and OpenAI admits in its own write-up that rival hardware may move a long way before Jalapeño is fully deployed. Jalapeño is less a win over Nvidia than an insurance policy on OpenAI's largest bill.
TechRepublic raised the question that matters to everyone downstream: OpenAI has not said whether customers will ever get to choose this hardware, or how the efficiency gains will show up in pricing. Cheaper inference for OpenAI is not automatically cheaper inference for you.
If you buy AI instead of building it
If you run a business in the Fox Valley, none of this is your problem to solve. You do not pick chips. But the cost and speed of inference sets the price of every AI-assisted phone answer or quote draft you might automate, and this week's news says that price keeps falling, just on a slower calendar than the posts suggest.
Two habits follow from that. Do not build a plan that only works if AI costs drop by half next quarter; build one that pays for itself at today's prices and improves as the bill shrinks. And treat vendor benchmarks the way SemiAnalysis treated OpenAI's: ask which comparison was chosen, what was skipped, and when it ships. That works just as well on a $200-a-month software pitch as on a chip.
If you want a second pair of eyes on where AI would pay for itself in your shop right now, New Face Design's free process audit is the place to start. We look at what you already do by hand and say plainly what is worth automating.