OpenAI's New AI Chip Could Challenge NVIDIA's Dominance in AI Computing

featured-image

OpenAI’s New AI Chip Could Challenge NVIDIA’s Dominance in AI Computing

OpenAI is taking a significant step beyond building artificial intelligence models. The company is now designing the computing hardware needed to run them, putting it on a potential collision course with NVIDIA in one of the most important parts of the AI infrastructure market.

OpenAI’s first custom AI processor, Jalapeño, was developed in partnership with Broadcom and is designed specifically for large language model inference—the computing process that turns a trained AI model into an actual response for users. OpenAI says early testing shows the chip can deliver both higher throughput and lower latency than the commercial systems used in its benchmark comparisons. (OpenAI)

The announcement does not mean NVIDIA’s dominance is about to disappear. NVIDIA remains deeply embedded in AI training and data-center infrastructure, and OpenAI itself continues to work with the chipmaker. But Jalapeño signals a broader change: the biggest AI companies increasingly want greater control over the hardware underneath their models.

OpenAI Is Building More of Its Own AI Stack

For years, the AI industry’s rapid expansion has depended heavily on NVIDIA GPUs. These processors became the standard hardware for training and running increasingly sophisticated AI models because of their computing power, software ecosystem and ability to operate at massive scale.

OpenAI has been one of the biggest beneficiaries—and customers—of that ecosystem.

But AI workloads are becoming enormous. Every ChatGPT question, coding request, image generation task and AI-agent interaction requires computing resources. As usage grows, even relatively small improvements in the cost and efficiency of each response can translate into substantial savings.

That is where custom silicon becomes attractive.

OpenAI unveiled Jalapeño in June as its first custom inference processor, developed with Broadcom. The company said the chip was designed from the ground up around the requirements of large language models and was developed from initial design to manufacturing tape-out in nine months. (OpenAI)

OpenAI and Broadcom had previously announced plans for a much larger custom-accelerator program targeting 10 gigawatts of AI infrastructure, with deployments expected to begin in the second half of 2026 and continue through 2029. (OpenAI)

That makes Jalapeño more than an isolated experiment. It is the first piece of a broader strategy to build a multi-generation computing platform around OpenAI’s models and products.

Why Inference Is Becoming So Important

The distinction between training and inference is central to understanding OpenAI’s chip strategy.

Training is the process of teaching an AI model by exposing it to enormous quantities of data. Inference happens afterward, when the trained model processes a user’s request and produces an answer.

As AI services become mainstream, inference represents an enormous and growing computing workload.

A model can be trained once, but it may serve millions or billions of requests afterward. For a company operating AI products at OpenAI’s scale, making each inference faster and more energy efficient can have a meaningful impact on operating costs.

OpenAI says its Jalapeño testing shows the chip can perform more AI work per unit of power while also reducing response latency. The company’s latest benchmark results were based on InferenceX testing involving models including GPT-OSS 120B, DeepSeek R1 and Kimi K2. (OpenAI)

That combination is particularly important for AI agents.

A traditional chatbot might answer a question with a handful of generated responses. An AI agent may perform dozens of reasoning steps, make multiple tool calls and interact with external systems before completing a task. Faster inference can therefore make an agent feel substantially more responsive while potentially reducing the computing cost of each task.

OpenAI Says Jalapeño Can Outperform NVIDIA Hardware

OpenAI’s latest benchmark announcement has attracted attention because it directly highlights performance against commercial NVIDIA systems.

According to OpenAI, Jalapeño achieved higher throughput per kilowatt and lower token latency than the commercial systems included in its InferenceX comparisons. The company says the results demonstrate that its architecture can avoid the traditional tradeoff between throughput and latency. (OpenAI)

Independent reporting on the company’s benchmark claims indicates that Jalapeño delivered between roughly 1.5 and 1.9 times more AI work per watt and 1.7 to 3.6 times lower end-to-end latency than NVIDIA’s GB200 and GB300 systems across several tested models. Those figures come from OpenAI’s cited benchmark results and should be viewed in the context of specific inference workloads rather than as proof that Jalapeño is universally faster than NVIDIA hardware. (The Verge)

That distinction matters.

A specialized accelerator does not need to beat a general-purpose AI platform at every task to be commercially valuable. If OpenAI can make its own chip significantly more efficient for the workloads it performs most frequently, it could shift a substantial portion of its computing demand away from conventional GPUs.

This Is Not Yet a Full NVIDIA Replacement

Despite the headlines, Jalapeño does not currently represent a complete alternative to NVIDIA’s AI computing platform.

The first-generation processor is focused on inference, rather than the broad range of workloads involved in training frontier AI models. OpenAI still relies on outside hardware, including NVIDIA technology, for significant portions of its computing infrastructure. (Axios)

NVIDIA also has advantages that extend far beyond the physical chip.

Its CUDA software ecosystem, networking technology, developer tools, systems architecture and enormous installed base create a substantial barrier for competitors. AI developers have spent years optimizing software for NVIDIA hardware, making the ecosystem itself an important competitive asset.

NVIDIA is also continuing to evolve its own infrastructure. The company is moving from its Blackwell generation toward Rubin while expanding beyond GPUs into CPUs, networking and complete AI systems. (Reuters)

So the more realistic interpretation is not that OpenAI is replacing NVIDIA overnight. Instead, OpenAI is attempting to reduce how dependent its future growth is on a single hardware supplier.

Custom Chips Could Change the Economics of AI

The biggest potential consequence of OpenAI’s strategy may not be a direct battle over chip benchmarks. It could be a change in the economics of running AI services.

Custom chips allow a company to optimize hardware for its specific workloads rather than paying for capabilities it may not use.

For OpenAI, that means designing around the characteristics of its own models, inference systems and future product roadmap. The company says its engineers used insights from its models, kernels, serving systems and product requirements when designing Jalapeño. (OpenAI)

That creates a potentially powerful feedback loop.

Better models can inform better hardware. Better hardware can reduce the cost of running those models. Lower costs can support more usage. More usage generates additional information about how people interact with AI, which can then influence future model and infrastructure development.

OpenAI has explicitly described this as a full-stack strategy spanning data centers, chips, models, developer products and consumer and enterprise services. (OpenAI)

If successful, that approach could give OpenAI more control over one of the largest expenses associated with operating AI at global scale.

OpenAI Is Part of a Much Bigger Shift

OpenAI is hardly alone in pursuing custom AI silicon.

Google has developed its Tensor Processing Units, while Amazon has built its own AI accelerators. Other major technology companies are also working with semiconductor specialists to develop chips tailored to their workloads.

Recent developments show how quickly this trend is expanding. Google, for example, has entered a major custom-AI-chip relationship with Marvell, illustrating the industry’s desire to diversify its hardware supply and optimize computing infrastructure for specific workloads. (Reuters)

The motivation is straightforward: demand for AI computing is growing so quickly that relying exclusively on commercially available accelerators can create enormous costs and supply constraints.

The result could be a more fragmented AI hardware market in which NVIDIA remains a dominant provider while hyperscalers and AI laboratories increasingly supplement GPUs with their own specialized processors.

NVIDIA Still Has Major Advantages

NVIDIA’s position should not be underestimated.

The company’s strength comes from the combination of hardware, software and infrastructure rather than from individual chip specifications alone. Developers know how to program for NVIDIA’s ecosystem, cloud providers have built infrastructure around it, and AI companies have invested heavily in optimizing models for its technology.

NVIDIA is also investing aggressively in the next phase of AI infrastructure.

The company’s financial and strategic position remains formidable, even as investors increasingly question whether the extraordinary growth of AI infrastructure spending can continue indefinitely. Reuters reported this week that NVIDIA is preparing for its next earnings report as the company transitions from Blackwell toward Rubin while facing increasing competition from AMD, Intel and custom chips developed by major technology companies. (Reuters)

OpenAI’s own relationship with NVIDIA further illustrates the complexity of the market. OpenAI has secured significant infrastructure commitments involving NVIDIA while simultaneously developing its own silicon. (OpenAI)

In other words, this is not necessarily a winner-takes-all contest.

OpenAI can use NVIDIA hardware where it makes sense while deploying custom processors where specialized economics provide an advantage.

The Bigger Battle Is Over Control of AI Infrastructure

The most important part of Jalapeño may therefore be what comes next.

OpenAI says the processor is the first generation of a multi-generation platform and expects initial deployment by the end of 2026, with the broader program designed to expand over subsequent generations. (OpenAI)

If future versions become capable of handling a wider range of AI workloads, OpenAI could gradually reduce its dependence on external accelerators. That would give the company more control over capacity, performance, energy consumption and infrastructure costs.

It could also force NVIDIA and other chipmakers to compete on more than raw computing power.

Efficiency, latency, networking, software compatibility and the ability to optimize chips for rapidly changing AI models could become equally important.

For NVIDIA, that means its biggest challenge may not come from one rival chip. It could come from the growing willingness of its largest customers to design around the GPU rather than simply buy more of it.

What OpenAI’s Chip Means for the AI Industry

Jalapeño is unlikely to overthrow NVIDIA’s position in AI computing in the near term. The chip is specialized, deployment is still beginning, and NVIDIA’s broader hardware and software ecosystem remains extremely difficult to replicate.

But the development is significant because OpenAI is no longer content to compete solely at the model layer.

The company is moving deeper into the infrastructure stack, following a path already taken by some of the world’s largest technology companies. If OpenAI can turn its early performance results into reliable, large-scale deployments, custom silicon could become an increasingly important part of how the company meets the enormous computing demands of ChatGPT, Codex and future AI agents.

The larger story is that the AI hardware market is entering a new phase. NVIDIA may remain the industry’s leading supplier, but the biggest AI companies increasingly want to control more of the machines that power their products.

And if that trend continues, the next battle in AI may be fought not only over who builds the smartest models, but over who controls the silicon those models run on.

0 comments
2

2 Comments

Micle harison

June 7, 2019

Lorem ipsum dolor sit amet, usu ut perfecto postulant deterruisset, libris causae volutpat at est, ius id modus laoreet urbanitas. Mel ei delenit dolores.

John Doe

June 7, 2019

Some consultants are employed indirectly by the client via a consultancy staffing company.

Leave a comment