Nvidia unveiled its next-generation Rubin AI platform at GTC 2026 on Monday, introducing six new chips that the company says will cut the cost of AI inference by a factor of ten compared to its current Blackwell architecture. CEO Jensen Huang announced the full platform during his keynote at the SAP Center in San Jose, where more than 30,000 attendees gathered from over 190 countries for what has become the largest annual AI conference in the world.
The Rubin platform includes the Vera CPU with 88 custom Olympus cores, the Rubin GPU with 50 petaflops of NVFP4 compute, a new NVLink 6 Switch delivering 3.6 terabytes per second of bandwidth per GPU, the ConnectX-9 SuperNIC, the BlueField-4 DPU, and the Spectrum-6 Ethernet Switch. Together, a full Vera Rubin NVL72 rack delivers 260 terabytes per second of total bandwidth.
What the Rubin Platform Means for AI Costs
The most significant number for founders and businesses building on AI: Nvidia claims Rubin will reduce token generation costs to roughly one-tenth of Blackwell’s current pricing. Training large mixture-of-experts models will require four times fewer GPUs. For startups spending heavily on cloud compute, the cost reduction could meaningfully change the economics of building an AI business.
“Rubin arrives at exactly the right moment, as AI computing demand for both training and inference is going through the roof,” Huang said during the keynote.
Major Cloud Providers and AI Labs Line Up
Nvidia announced that Rubin is already in full production and will be available from partners in the second half of 2026. AWS, Google Cloud, Microsoft, and Oracle Cloud Infrastructure will be among the first to deploy Vera Rubin instances, alongside cloud partners CoreWeave, Lambda, Nebius, and Nscale.
The company also revealed a multiyear strategic partnership with Thinking Machines Lab, the AI startup founded by former OpenAI CTO Mira Murati, to deploy at least one gigawatt of Vera Rubin systems for frontier model training. That partnership puts Thinking Machines in the same tier as the largest AI infrastructure buyers in the world.
Among the AI labs and startups planning to adopt Rubin are Anthropic, OpenAI, Meta, xAI, Cohere, Mistral AI, Perplexity, Cursor, Harvey, and Runway. Anthropic CEO Dario Amodei called the platform a step forward, noting that the efficiency gains “enable longer memory, better reasoning and more reliable outputs.”
GTC 2026 Signals the Scale of What Is Coming
The conference, running March 16 through 19, also featured the launch of OpenClaw, which Nvidia described as the fastest-growing open source project in history. OpenClaw is a framework for building local AI agents that run on personal hardware, including Nvidia’s DGX Spark system, without requiring cloud connectivity. The company also announced that pharmaceutical giant Lilly launched what it called the most powerful AI factory wholly owned by a drug company.
For the broader startup ecosystem heading into 2026, the Rubin platform represents a generational shift in AI infrastructure. If Nvidia delivers on the 10x cost reduction it is promising, the barrier to building and deploying AI products will drop significantly, opening the door for smaller companies to compete with well-funded labs on inference quality and speed.
Rubin-based products are expected to ship from hardware partners beginning in the second half of this year.



