NEWS

Amazon and Cerebras Partner to Build the Fastest AI Inference Service in the Cloud

Data center servers powering cloud AI inference for Amazon AWS and Cerebras partnership

Amazon Web Services and Cerebras Systems announced a multiyear partnership that pairs AWS’s proprietary Trainium chips with Cerebras’s massive Wafer-Scale Engine to deliver what both companies say will be the fastest AI inference available in the cloud. The new service will be accessible through Amazon Bedrock starting in the second half of 2026.

The deal represents the most significant challenge yet to Nvidia’s dominance in AI computing infrastructure. Rather than relying on a single chip architecture, the two companies are splitting AI inference into two specialized stages, an approach known as disaggregated inference.

How the Partnership Works

Under the arrangement, AWS Trainium 3 chips handle the prefill stage, which processes and makes sense of user prompts. Cerebras’s WSE-3 chip then takes over for the decode stage, generating the actual output. The two systems communicate through AWS Elastic Fabric Adapter networking.

The Cerebras WSE-3 is no ordinary processor. At 56 times the size of the largest GPU on the market, the chip packs 900,000 cores and 44 gigabytes of on-chip SRAM. It ships inside a water-cooled appliance called the CS-3, roughly the size of a mini-fridge. Its memory bandwidth reaches 27 petabytes per second, which Cerebras says is more than 200 times faster than Nvidia’s latest networking technology.

“The result will be inference that’s an order of magnitude faster and higher performance than what’s available today,” said David Brown, Vice President of Compute and ML Services at AWS.

Why It Matters for AI Startups

For founders building AI-powered products, inference speed directly affects user experience. Chatbots, coding assistants, and real-time AI applications all depend on how quickly a model can generate responses. The partnership promises to increase output generation speed by a factor of five compared to current cloud offerings.

The service will launch first in the US-East (Northern Virginia) and US-West (Oregon) AWS regions by the third quarter of 2026. AWS plans to offer leading open-source large language models and its own Amazon Nova models on the Cerebras hardware.

“Every enterprise around the world will be able to benefit from blisteringly fast inference within their existing AWS environment,” said Andrew Feldman, Founder and CEO of Cerebras.

A Growing List of High-Profile Customers

Cerebras has been building momentum ahead of a planned initial public offering expected in the second quarter of 2026. The startup already counts OpenAI among its customers through a deal reportedly worth more than $10 billion for 750 megawatts of computing infrastructure through 2028. Cognition and Mistral are also using Cerebras technology.

The Amazon partnership adds the world’s largest cloud provider to that roster, giving Cerebras significant credibility as it prepares to go public. For AWS, the deal provides a differentiated AI offering that does not depend entirely on Nvidia hardware, which remains in high demand and subject to ongoing export policy debates.

The Bigger Picture

The partnership arrives as demand for AI computing continues to outpace supply. Major cloud providers have been investing tens of billions of dollars in data centers, custom chips, and infrastructure deals to keep up with enterprise AI adoption. Google recently committed $1 billion to expand its data center in North Carolina, and Meta announced four new generations of its in-house MTIA chips earlier this week.

Financial terms of the Amazon-Cerebras arrangement were not disclosed. AWS did not specify pricing for the new inference service, though it will be available through the existing Bedrock platform that enterprise customers already use to access foundation models.

Read More From the NEWS desk