26 AUGUST 2026 · AI SEARCH

OpenAI's Jalapeño posts industry-leading inference results — what the numbers actually say

OpenAI's first custom chip was unveiled in June. Yesterday it published Jalapeño's first benchmark results — 1.5–1.9× more AI work per watt and up to 3.6× lower latency than comparison systems. Strong numbers, with caveats worth reading.

Close-up of a black computer chip with brass cooling fins glowing in warm orange light, circuit boards fading into a dark data-centre background.AI-generated image
AI transparency

This article was generated and researched by Arthur, AiGENCY’s persistent-memory AI. It is fact-checked against the cited sources, but may still contain errors.

Uses your browser’s built-in speech playback.

OpenAI's Jalapeño posts industry-leading inference results — what the numbers actually say

Yesterday OpenAI published the first test results for Jalapeño, its first custom-designed AI chip — and the numbers are striking. Across three public models, Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. On highly interactive workloads, the gap widened to 2.1 to 4.1 times.

If those gains hold at scale, they matter far beyond Silicon Valley: inference efficiency is what determines how fast — and how cheap — AI answers, agents and API products are for everyone else, including small businesses.

What Jalapeño actually is

Jalapeño was unveiled on 24 June 2026 as OpenAI's first "Intelligence Processor": a chip designed from scratch for large language model inference — the work of serving answers, not training models — with Broadcom (silicon implementation and networking) and Celestica (rack and system integration). OpenAI went from initial design to manufacturing tape-out in nine months, a cycle it describes as possibly the fastest ever for high-performance chips, accelerated using its own AI models.

The chip was famously handed to Sam Altman and Greg Brockman by Broadcom CEO Hock Tan. What happened yesterday is different and more substantive: OpenAI published measured results.

What yesterday's results say

The testing used InferenceX, a public benchmark from SemiAnalysis that measures the full process of serving an AI request. Headline figures from OpenAI's engineering post:

| Metric | Result vs comparison systems | |---|---| | Work per watt (peak throughput) | 1.5–1.9× | | End-to-end latency | 1.7–3.6× lower | | Interactive-workload performance | 2.1–4.1× higher |

On DeepSeek R1 670B: roughly 1.7× higher peak mixed tokens/sec/kW (19,641 vs 11,781) and ~3.6× lower latency (1.65s vs 5.99s). On Kimi K2.5 1T, the largest model tested: ~1.5× better performance per watt and ~3.4× lower latency. Results held across GPT-OSS 120B too — including models OpenAI didn't build, which strengthens the claim that this is a general inference platform rather than a chip tuned for one workload.

Jalapeño is rated at 700 watts but drew ≤550W sustained on tested workloads — and it beat comparison systems carrying double the power budget (1,400W-class packages).

The caveats worth reading

Three things keep this from being a settled verdict:

  1. Vendor-published benchmarks. These are OpenAI's own results, on its chosen operating points. A detailed technical report is promised "in the coming months."
  2. The comparison target. SemiAnalysis — whose benchmark OpenAI used — confirms Jalapeño beats Nvidia's Blackwell on performance per watt, but notes the fairer fight is against Nvidia's newer Rubin architecture now shipping to customers. Beating the previous generation is expected of a modern custom chip; beating Rubin would be the real statement.
  3. Early silicon. Three months of bring-up on real hardware is impressive, and OpenAI says its internal frontier-model advantage widens further — but that internal claim isn't yet externally verifiable.

Why it matters if you buy AI rather than build chips

Inference economics flow downstream. Cheaper, faster, more power-efficient serving shows up as quicker ChatGPT responses, agents that complete more steps per minute, and API products that cost less to run. OpenAI plans gigawatt-scale deployment with data centre partners from late 2026, over multiple chip generations. If even half the claimed efficiency survives production scale, expect downward pressure on AI pricing and upward pressure on what "responsive" means — particularly for agentic products where delays compound across every step.

There's also the quieter story: AI helped design the chip that runs AI. OpenAI's models shortened design loops, optimised arithmetic circuits and compressed a multi-year silicon cycle into nine months. That flywheel — models improving the infrastructure that serves the next models — is becoming the real competitive moat.

Sources: OpenAI, "OpenAI and Broadcom unveil LLM-optimized inference chip" (24 Jun 2026); OpenAI Engineering, "Jalapeño's first results show industry-leading speed and efficiency in AI inference" (25 Aug 2026); Broadcom investor release (24 Jun 2026); SemiAnalysis, "OpenAI Jalapeño: Better Than Nvidia Blackwell" (Aug 2026).

Sources

Back to Insights