Pumpkin AI Studio — Imagination, rendered.

Pumpkin AI / Global intelligence

Start a brief
Menu

OpenAI's Jalapeño chip: the AI race beneath the model

OpenAI has published the first measured results for its custom Jalapeño inference chip. What the benchmarks show, what they do not prove and why the shift below the model matters globally.

A monumental black custom-compute structure threaded with restrained amber energyPumpkin frame / 01
Visual note

The AI race is moving beneath the model into an engineered stack of silicon, networks, cooling and power.

An engineer walking through a dark data centre of cooling pipes, power routes and compute racksPumpkin frame / 02
Context image

Efficiency per task matters, but the complete system—and the total demand placed on it—determines the real energy consequence.

01

What OpenAI actually announced

OpenAI's 25 August engineering report presents the first measured performance results for Jalapeño, its custom inference accelerator. The company tested the chip on GPT-OSS 120B, DeepSeek R1 and Kimi K2.5 using InferenceX, the public inference-benchmark framework maintained by SemiAnalysis. OpenAI reports 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems in those tests. For selected interactive workloads, it reports a 2.1 to 4.1 times performance advantage.

Those numbers are meaningful because inference is the part of AI that happens after a model has been trained: every answer, generated frame, agent action or software request has to be served somewhere. Lower latency can make an experience feel more immediate; greater throughput per watt can let the same electrical and cooling envelope handle more useful work.

The announcement is not a retail launch. OpenAI says Jalapeño is still going through production qualification and software validation. It plans to begin deploying the first generation in its own compute infrastructure by the end of 2026. No public sale, general cloud availability, customer price or confirmed external deployment date has been announced.

02

The benchmark is evidence, not a verdict

OpenAI's use of a public framework is a positive step because the workload definitions and measurement approach can be inspected. InferenceX is published under an Apache 2.0 licence and is designed to compare the full serving system rather than a single laboratory operation. That makes latency, throughput and power visible together instead of treating speed as the only outcome.

The result still needs a careful label: it is company-reported performance for hardware that outside users cannot yet obtain. A benchmark publisher can standardise a test, but broad independent replication requires other teams to run the same chip, software and configuration. That is not currently possible at market scale. Software stacks also improve, so any hardware comparison is a point-in-time picture rather than a permanent ranking.

OpenAI rates Jalapeño at 700 watts and says the chip stayed at or below 550 watts during the published workloads. That describes the accelerator measurements in these tests; it is not the total electricity or environmental footprint of a finished AI service. Networking, memory, host systems, cooling, data-centre overhead and the number of user requests all matter. Responsible reporting keeps those system boundaries visible.

  • Was quality held constant while latency and throughput were compared?
  • Is the result independently repeatable on hardware available outside the vendor's laboratory?
  • Does the power figure cover one accelerator, the server, the rack or the complete facility?
  • Will measured efficiency improve customer price, responsiveness or availability?
03

Why inference has become the strategic layer

Training creates a model; inference turns that model into a service. As AI adoption expands, the economics of serving can become as important as the cost of the original training run. A popular assistant may answer millions of requests. A creative system may generate many alternatives. An agent may make a long chain of model calls before completing one visible task. Small delays and energy costs can compound across every step.

That changes what counts as a competitive advantage. A better model can attract demand, but a better serving system can determine whether the product is fast enough, affordable enough and available in enough regions to retain it. Custom silicon gives a company another way to tune that system around the kinds of models and interactions it expects to operate most often.

It also explains why a chip story is a product story. If infrastructure reduces waiting time, the interface feels more capable even when the underlying model has not changed. If it improves useful work per watt, a provider may be able to serve more demand within a constrained power envelope. Neither benefit is automatic for customers, but both can shape the next generation of AI services.

04

Custom silicon does not mean a closed one-chip future

Jalapeño is the product of a partner system, not a solitary piece of engineering. OpenAI says it designed the architecture, Broadcom handled silicon implementation and networking, and Celestica contributed board, rack and system design. The chip is one part of a broader stack in which models, compilers, serving software, interconnects and facilities have to work together.

OpenAI's wider infrastructure statement is equally important. The company says it will continue using accelerators and services from partners including NVIDIA, AMD, Microsoft, AWS, CoreWeave and others. The reasonable interpretation is therefore diversification and leverage, not the immediate replacement of every general-purpose accelerator with a single internal design.

This pattern is global. The largest AI platforms increasingly want control over more of the path from electricity to user experience, while also retaining multiple suppliers. For governments and businesses, that can create both resilience and concentration risk: more technical options inside a small number of very large platforms, but greater capital and infrastructure barriers for everyone trying to compete with them.

A precise illuminated form inside a dark angular enclosure03
A benchmark isolates a measurable result; real-world value depends on the complete system around it.
05

Efficiency matters because energy demand is still rising

The timing of the announcement is not accidental. In April 2026, the International Energy Agency reported that global data-centre electricity use had risen 17% in 2025. It also found that electricity use per AI task was falling rapidly while overall demand kept increasing as more people used AI and more intensive uses, including agents, expanded. The IEA expects data-centre electricity consumption to double by 2030 and power use at AI-focused facilities to triple.

That is the central tension in performance-per-watt claims. More efficient hardware can reduce the electricity required for one task, but cheaper and faster inference can also unlock more tasks. Total demand can therefore rise even as the unit footprint falls. A credible sustainability claim has to report both: efficiency per useful outcome and the absolute energy, water and infrastructure required at scale.

The same IEA analysis points to physical bottlenecks around transformers, advanced chips, generation equipment and grid connections. For AI platforms, infrastructure efficiency is becoming a way to stretch scarce capacity. For communities and policymakers, it does not remove the need to ask where new facilities are built, how they are powered and who pays for the supporting grid.

06

What businesses and creative teams should watch next

Most organisations do not need to choose an accelerator. They do need to understand what infrastructure changes can alter in the products they buy. The useful signals will be observable: faster response under load, fewer capacity limits, wider regional access, clearer service guarantees and pricing that reflects real efficiency rather than a benchmark headline.

Procurement teams should ask providers which performance claims apply to the exact model, quality setting and workload they intend to use. Creative and media teams should measure the waiting time across a complete human task, not only the generation time of one output. A nominally faster model can still create a slower workflow when queuing, retries, asset transfer or review become the bottleneck.

There is also a strategic question. If the best models, chips and delivery systems are increasingly integrated inside a few platforms, businesses should preserve portability where it matters: retain their own approved assets and outputs, document quality requirements, avoid unnecessary dependence on one proprietary interface and keep human judgement independent of the underlying supplier.

  • Availability: Is the capability deployed now, in which regions and for which customers?
  • Economics: Is lower serving cost visible in the contract or only in the provider's margin?
  • Reliability: Does the system sustain its result under real concurrent demand?
  • Measurement: Are latency, quality, energy and total task completion evaluated together?
  • Resilience: Can important work move if one model, platform or region becomes unavailable?
07

The Pumpkin AI conclusion: the stack is now part of the intelligence

Jalapeño matters less as a claim that one chip has won and more as evidence that the AI race has entered a full-stack phase. Models remain visible, but chips, networking, power, serving software and operations increasingly decide how much intelligence reaches people, how quickly it arrives and what it costs to sustain.

Pumpkin AI will judge this shift through delivered outcomes rather than corporate scale alone. The right questions are not only how fast a benchmark ran, but whether the technology is available, independently testable, responsibly powered and useful in the hands of real organisations. The companies that can answer all four will shape the next phase of AI more convincingly than the companies with the loudest specification sheet.

FAQ

Questions worth asking.

What is OpenAI's Jalapeño chip?

Jalapeño is a custom accelerator designed by OpenAI for AI inference, with Broadcom responsible for silicon implementation and networking and Celestica contributing board, rack and system design. It is intended to run trained models efficiently when they serve user requests.

Can businesses buy or use Jalapeño now?

No public sale or general cloud availability has been announced. OpenAI says production qualification and software validation are continuing, with initial deployment planned inside its own compute infrastructure by the end of 2026.

Does Jalapeño replace NVIDIA GPUs?

Not according to OpenAI's published strategy. The company says it will continue deploying accelerators and infrastructure from multiple partners, including NVIDIA, while adding its own custom silicon. Jalapeño is better understood as an additional specialised option than an immediate universal replacement.

Why does AI performance per watt matter?

It measures how much useful model serving can be completed within an electrical limit. Better performance per watt can increase capacity and potentially reduce the unit cost of inference, although total energy demand may still rise as AI use expands.

Are the Jalapeño benchmark results independently verified?

The tests use the public InferenceX framework, which improves transparency, but the published results come from OpenAI and the chip is not broadly available for outside replication. They should be treated as measured company results rather than a final independent market ranking.

Sources

Sources and further reading.

Continue

Need to act on the signal?

Turn the shift into a useful creative decision.

Talk to Pumpkin AI