OpenAI and Broadcom Unveil Jalapeno
OpenAI Jalapeno

OpenAI⁠ has officially announced Jalapeño, its first custom-built AI inference chip, developed in partnership with  Broadcom⁠. The new processor is purpose-built for large language model (LLM) inference, the compute-heavy process behind generating responses in products like ChatGPT, Codex, and API-powered AI applications.  

This marks a major strategic shift for OpenAI. Until now, the company relied heavily on GPUs from  NVIDIA⁠ and other infrastructure partners to power its AI workloads. With Jalapeño, OpenAI is moving deeper into custom silicon, aiming to gain tighter control over performance, power efficiency, and infrastructure costs.  

What Is Jalapeño?

OpenAI describes Jalapeño as an “Intelligence Processor”—an AI accelerator specifically optimized for inference rather than model training. Instead of focusing on building massive training clusters, the chip is designed to serve millions of real-time AI requests more efficiently.

According to OpenAI, Jalapeño is architected around the unique workload patterns of LLM inference, where low latency, memory bandwidth, and efficient token generation are critical.  

Key focus areas include:

  • Faster AI inference

  • Lower latency for responses

  • Better performance per watt

  • Reduced infrastructure cost

  • Higher throughput for AI serving

This makes the chip particularly valuable for high-scale consumer products like ChatGPT and enterprise AI APIs.

Built With Broadcom

Broadcom played a central role in designing and engineering the chip. The company has extensive experience in custom AI silicon, including work on accelerator architectures for hyperscalers and cloud providers.

Broadcom CEO Hock Tan said Jalapeño performs competitively with top-tier AI accelerators such as NVIDIA Blackwell and Google TPUs for inference workloads.  

Manufacturing will reportedly be handled by  TSMC⁠, while server integration is being supported by  Celestica⁠.  

Why OpenAI Built Its Own Chip

The AI industry is facing a major compute bottleneck.

Running frontier AI models requires enormous amounts of expensive hardware, and access to GPUs—especially from NVIDIA—has become one of the biggest constraints in AI scaling.

By developing custom chips, OpenAI can:

  • Reduce dependency on third-party GPU supply

  • Lower inference costs

  • Optimize hardware for its own models

  • Improve service reliability at scale

Already Running in OpenAI Labs

OpenAI says engineering samples of Jalapeño are already operational in internal labs and have been tested on production-like AI workloads, including GPT-5.3 Codex Spark inference scenarios. Commercial deployment is targeted by late 2026.  

The company also confirmed that Jalapeño is only the first chip in a multi-generation hardware roadmap, suggesting OpenAI is building a long-term full-stack compute platform.