OpenAI Deploys GPT-5.3-Codex-Spark on Cerebras Chips, Boosts Coding Speed

OpenAI has launched its GPT-5.3-Codex-Spark coding model on Cerebras hardware, marking its first production deployment on non-Nvidia chips. The model delivers code at more than 1,000 tokens per second, significantly outpacing its predecessors and competitors. API access is being rolled out selectively, with initial availability to ChatGPT Pro subscribers.

ShareShare

OpenAI has introduced its first production artificial intelligence model to run on non-Nvidia hardware: the GPT-5.3-Codex-Spark coding model powered by Cerebras chips. The launch, announced on Thursday, represents a notable shift in AI infrastructure, as the industry has largely relied on Nvidia graphic processing units (GPUs) for training and deploying large language models (LLMs).

GPT-5.3-Codex-Spark is engineered specifically for code generation and software engineering tasks. In initial production benchmarks, the model is able to deliver code at over 1,000 tokens per second—a unit that refers to portions of data processed in sequence. This marks an estimated 15-fold improvement in generation speed compared to previous iterations. For context, Anthropic's Claude Opus 4.6, a rival coding assistant model, delivers about 2.5 times its previous speed, but still lags significantly behind Spark’s benchmark. While Claude is a larger and more sophisticated model, Spark’s execution speed is currently a key differentiator.

The migration from Nvidia’s dominant GPUs to the custom-designed chips from Cerebras signals an evolving landscape in AI hardware. Cerebras systems use large, plate-sized processor wafers to handle the massive computations required for training and serving contemporary LLMs. This collaboration demonstrates increasing willingness by major AI companies to diversify their hardware partners for performance and supply chain resilience.

“Cerebras has been a great engineering partner, and we're excited about adding fast inference as a new platform capability,” commented Sachin Katti, head of compute at OpenAI. Inference speed refers to how quickly an AI model processes and generates output once it has been trained.

Initially, Codex-Spark is being offered as a research preview, accessible to ChatGPT Pro subscribers at a subscription price of $200 per month. The model can be used via the dedicated Codex app, a command-line interface, and an extension for Visual Studio Code, a popular code editor. OpenAI intends to extend API access to select design partners over time.

Codex-Spark ships with a 128,000-token context window at launch. The context window indicates how much preceding data the model can consider when generating new responses, a vital factor for handling complex coding tasks. Currently, the model supports only text input and output.

The move to embrace Cerebras hardware underscores a larger trend among AI developers to seek alternatives to Nvidia for greater cost efficiency, improved energy performance, and more robust global supply chains. While much of the industry remains reliant on Nvidia’s GPUs, this deployment could spur growing adoption and development of specialized AI chips tailored to specific workloads.

For software developers and enterprises, faster code generation models could accelerate product development cycles, streamline coding assistance tools, and lower infrastructure costs associated with AI-powered programming.

As leading AI companies continue to test and adopt new hardware platforms, the competitive dynamics in both the AI model and chip markets are likely to intensify, potentially reshaping how advanced neural networks are deployed at scale.

Source: arstechnica.com.

Related Posts

CoreWeave Prioritises Speed by Leasing UK Data Centre Space for AI

CoreWeave, a US cloud computing firm backed by Nvidia, is leasing data centre space in the UK to accelerate the rollout of AI infrastructure. The strategy aims to address soaring demand more rapidly than building new facilities from scratch.

Microsoft Uses Agentic AI to Advance Majorana 2 Quantum Chip

Microsoft has introduced its Majorana 2 quantum chip, achieving a significant leap in qubit reliability. The chip's development serves as a demonstration of the company's Discovery agentic AI platform, now available for enterprise and research use.

Alphabet Seeks $80 Billion to Fund Expansion of AI Capabilities

Alphabet has announced plans to raise $80 billion to finance further development of its artificial intelligence infrastructure. The capital will be used to enhance computing power and support AI research, reflecting growing investment demands across the technology sector.

The Essential Weekly Update

Stay informed with curated insights delivered weekly to your inbox.