OpenAI partners with Cerebras

SAN FRANCISCO — January 2026 — OpenAI has announced a new partnership with Cerebras aimed at significantly expanding its AI inference capacity using Cerebras’ wafer-scale computing systems.

Under the agreement, OpenAI will leverage Cerebras’ specialized hardware to run parts of its AI workloads, particularly inference tasks that require extremely fast response times. The move is designed to complement OpenAI’s existing cloud infrastructure partnerships and help meet growing global demand for AI-powered services.

OpenAI said the collaboration allows it to diversify compute resources while improving performance for certain classes of models and workloads.

The partnership centers on Cerebras’ wafer-scale engines, which integrate an entire silicon wafer into a single processor. This architecture is optimized for large AI models, enabling faster inference, simplified scaling, and reduced system complexity compared with traditional GPU-based clusters.

According to the companies, Cerebras systems can deliver near-instant responses for large language models, making them well suited for real-time and high-throughput applications.

OpenAI emphasized that the partnership does not replace its existing relationships with cloud providers. Instead, Cerebras adds another layer to its compute strategy, giving OpenAI more flexibility to match specific workloads with the most effective hardware.

The collaboration reflects a broader trend in AI infrastructure toward heterogeneous compute, where specialized chips are used alongside general-purpose accelerators to improve efficiency, resilience, and performance at scale.

Cerebras said working with OpenAI validates its approach to AI hardware and highlights the growing need for alternative architectures as model sizes and usage continue to increase.

The partnership is live, with Cerebras systems already being used to support OpenAI workloads, and further expansion planned as demand grows.

Source: OpenAI