HUSTLE · GROW

Jensen Huang Computex 2026 Keynote Recap for Founders

AI data center infrastructure representing NVIDIA's Computex 2026 keynote vision for agentic computing
0:00
0:00🎧 19 min

On June 1, 2026, Jensen Huang walked onto a stage in Taipei wearing his usual leather jacket and spent two hours making a single argument. NVIDIA isn’t a chip company anymore. It’s not even a systems company. “We’ve really become an infrastructure company,” Huang told the audience at GTC Taipei during COMPUTEX 2026. “Not just a GPU company, not just a systems company, but an infrastructure company to help you generate the maximum revenues, the maximum profit and to get there as soon as possible.”

Most tech reporters focused on the product announcements: RTX Spark, Nemotron 3 Ultra, Cosmos 3. Fair enough. But for founders building with AI, the real story wasn’t any single product. It was the aggregate signal. NVIDIA, the company that posted $215.9 billion in fiscal 2026 revenue and $120.1 billion in net income, just told the world it wants to be the infrastructure layer underneath every AI-native business. That changes the math on how you build, what you spend on cloud inference, and when local AI becomes a real option for your team.

The NVIDIA Computex 2026 keynote was Jensen Huang’s declaration that NVIDIA has transformed from a GPU maker into a full-stack AI infrastructure provider, unveiling the RTX Spark superchip for local AI agents, confirming Vera CPU production for hyperscale data centers, and launching Nemotron 3 Ultra as a 500-billion-parameter open model.

Last updated: June 2026

Quick answers

What did Jensen Huang announce at Computex 2026?

Jensen Huang used his two-hour COMPUTEX 2026 keynote on June 1 to declare NVIDIA an “infrastructure company,” unveil the RTX Spark superchip for local AI on laptops, confirm Vera CPU production for hyperscale agentic workloads, launch the 500-billion-parameter Nemotron 3 Ultra open model, introduce Cosmos 3 for robotics, and announce $150 billion in annual Taiwan investment.

What is NVIDIA RTX Spark?

RTX Spark is NVIDIA’s new superchip for Windows laptops and desktops, combining a 20-core Arm-based Grace CPU (co-developed with MediaTek) and a Blackwell GPU with 6,144 CUDA cores on a single TSMC 3nm package. It supports up to 128GB unified LPDDR5X memory and can run 120-billion-parameter AI models locally with million-token context windows. Laptops ship fall 2026.

What does it mean that NVIDIA is an infrastructure company?

NVIDIA is positioning itself as the full-stack layer underneath AI businesses, not just a GPU seller. The company now provides end-to-end infrastructure spanning chips (Vera CPU, Blackwell GPU), networking (NVLink), software (CUDA, NIM microservices), and open models (Nemotron 3). For founders, this means NVIDIA wants to be to AI what AWS became to cloud computing.

What did Jensen Huang announce at Computex 2026?

Huang’s keynote covered six major announcements, each reinforcing the same thesis: the AI industry has moved from training models to running agents, and NVIDIA built the infrastructure stack for that shift.

The headline product was RTX Spark, a new Windows-on-Arm superchip that Huang called “reinventing the PC.” It pairs a 20-core Grace CPU, co-developed with MediaTek, with a Blackwell-architecture GPU packing 6,144 CUDA cores and up to 128GB of unified LPDDR5X memory. All of it sits on a single TSMC 3nm package. NVIDIA says it can run 120-billion-parameter models and handle context lengths of one million tokens entirely on-device. Laptops from Dell, HP, ASUS, Lenovo, MSI, and Microsoft Surface will ship with RTX Spark in fall 2026.

Then came the confirmation founders who use cloud AI services have been waiting for: Vera CPU is in full production. Not on a roadmap. Shipping. Anthropic, OpenAI, xAI, Dell, Oracle, and CoreWeave are already running production agentic AI workloads on Vera. Huang called it “the first CPU designed for the agentic future,” built to coordinate AI models, storage systems, and large-scale compute clusters with twice the efficiency and 50% faster throughput than traditional rack-scale CPUs, according to NVIDIA’s official announcement.

The third pillar was Nemotron 3 Ultra, NVIDIA’s open AI model. At 500 to 550 billion parameters using a mixture-of-experts (MoE) architecture, it activates only about 50 billion parameters per token. That means 5x higher throughput and roughly 30% lower cost than competing models, per NVIDIA’s newsroom. It’ll be available on Hugging Face, ModelScope, OpenRouter, and as an NVIDIA NIM microservice. For founders who’ve been paying OpenAI or Anthropic API costs, an open model at this scale from the company that makes the hardware is a pricing pressure signal worth watching.

Server room with illuminated racks representing AI data center infrastructure like the kind NVIDIA builds

Cosmos 3 was the robotics play. NVIDIA launched what it calls the world’s first open Physical AI omnimodel, handling text, images, video, ambient sound, and robotic actions in a single inference pass. It ships in two sizes: Super at 32 billion parameters and Nano at 8 billion for edge deployment. Samsung, LG Electronics, Li Auto, and Doosan Robotics are confirmed adopters. For founders building anything that touches the physical world, from warehouse automation to delivery logistics, Cosmos 3 is NVIDIA’s bet that simulation-trained robots will run on its stack.

The geopolitical statement came last. Huang committed to $150 billion per year in Taiwan investment, a number roughly 10x NVIDIA’s previous annual spend in the country. The company broke ground on Constellation, a new Taipei campus that will house about 4,000 employees in the Beitou-Shilin Technology Park by 2030. “Taiwan is the epicenter of the AI revolution,” Huang said. The subtext: NVIDIA is tying its supply chain even more tightly to TSMC, the foundry that manufactures its most advanced chips.

He also teased something for the second half of 2026: “Grace Blackwell, Vera Rubin, and a surprise new product that hasn’t been announced yet.” No details. Just enough to keep the market guessing.

What does it mean when NVIDIA calls itself an infrastructure company?

NVIDIA generated $197.3 billion in data center revenue in fiscal 2026, up from $115.2 billion the year before, per its SEC filing. Data center now accounts for 91% of total revenue. The GPU gaming company that old-school investors remember? It’s a rounding error now.

When Huang says “infrastructure company,” he’s describing a shift that’s already happened in the numbers. NVIDIA doesn’t just sell chips to hyperscalers. It sells complete systems: GPUs, CPUs, networking (NVLink), software (CUDA, TensorRT, NIM microservices), and now open models. The pitch to enterprises isn’t “buy our GPU.” It’s “plug into our stack and we’ll help you generate maximum revenue from AI.”

For founders, this framing matters because it signals where NVIDIA’s incentives point. An infrastructure company needs customers who build on its platform, not customers who buy a chip once and leave. That means NVIDIA is motivated to make its platform accessible, interoperable, and cost-effective for builders at every scale. The RTX Spark announcement is the clearest proof: Huang wants NVIDIA’s AI stack on every laptop, not just in every data center.

The comparison NVIDIA wants you to make is AWS. Amazon Web Services didn’t just sell servers. It sold infrastructure that startups plugged into, and the developer community that grew on top became the moat. Huang’s Computex keynote was his version of that pitch: NVIDIA provides the AI factory, you build the applications.

There’s a catch for founders to watch, though. Infrastructure companies create vendor dependency. CUDA, NVIDIA’s programming framework, already dominates AI development. If RTX Spark laptops become the standard for running local AI, that tightens the lock-in further. The same CUDA binary that runs on an H100 data center GPU runs on RTX Spark without recompilation. That’s convenient if you’re already building on NVIDIA. It’s a competitive moat if you’re trying to leave.

How RTX Spark changes the cost math for AI startups

The economics of running AI are about to split into two tracks, and founders need to understand both.

Track one is cloud inference getting cheaper. Vera Rubin systems coming online throughout 2026 mean hyperscalers like AWS, Google Cloud, and Azure will have more compute capacity to sell, which pushes API prices down. That benefits every startup paying per-token costs today.

Track two is local inference becoming genuinely capable for the first time on laptop-class hardware. RTX Spark supports 120-billion-parameter models with million-token context windows running entirely on-device. That’s not a toy demo. A 120B model is roughly the size of GPT-3.5-class systems. Running it locally means zero API costs, zero latency, and no data leaving your device.

For a bootstrapped founder running a one-person AI business, this math is concrete. If you’re currently spending $200 to $500 per month on OpenAI or Anthropic API calls for internal tools, customer support bots, or content generation, an RTX Spark laptop that runs equivalent models locally eliminates that recurring cost entirely. The hardware is a one-time purchase. The inference is free after that.

That doesn’t mean cloud AI goes away. Models larger than 120B parameters, frontier reasoning tasks, and workloads that need to scale instantly still need cloud infrastructure. But the threshold for “what can run locally” just moved significantly. NVIDIA says the same CUDA code that runs on H100 data center hardware runs on RTX Spark without changes. If your AI development stack uses PyTorch with CUDA, llama.cpp, Flash Attention, or TensorRT, it works out of the box.

The competitive pressure this creates is real. Intel closed down roughly 4.7% on June 2 after the RTX Spark announcement, while NVIDIA gained more than 6%. AMD and Qualcomm, both competing in the AI PC space, now have to match NVIDIA’s combination of a custom Arm CPU, Blackwell GPU, and 128GB unified memory on a single package. Apple’s M-series chips handle local AI well, but they run on macOS. RTX Spark brings equivalent or superior local AI capability to the Windows platform, which is where most enterprise software still lives.

Vera CPU ships to the companies building your AI tools

If RTX Spark is the consumer-facing story, Vera CPU is the infrastructure story that affects founders indirectly but powerfully.

Vera is now in production at the companies that provide your AI backend. Anthropic uses it. OpenAI uses it. xAI uses it. When Huang said Vera is “the first CPU designed for the agentic future,” he meant it’s built specifically for the kind of AI workloads that run continuously, make decisions autonomously, and coordinate across multiple models and data sources. That’s the direction every major AI provider is heading: from chatbots you prompt to agents that work on your behalf.

For founders, the downstream effect is straightforward. More efficient infrastructure at the hyperscaler level means lower costs passed through to API customers. Vera’s claimed 2x efficiency improvement and 50% speed advantage over traditional CPUs, if it holds at scale, should translate into cheaper inference pricing from Anthropic, OpenAI, and the other providers running on it. The timeline for feeling that in your monthly API bill is probably 6 to 12 months as these providers negotiate new pricing based on lower infrastructure costs.

The agentic angle is the one to watch most closely. Vera was designed for workloads where AI agents run continuously, not in request-response bursts. If your product roadmap includes AI agents that monitor, decide, and act without human prompting, the infrastructure underneath those agents just got purpose-built hardware. That’s a signal to accelerate agent development timelines, not wait.

What is Nemotron 3 Ultra and why should founders care?

Nemotron 3 Ultra is NVIDIA’s open-weights AI model, and it’s big enough to matter in the open-source AI race.

At 500 to 550 billion parameters, it uses a latent mixture-of-experts architecture that activates only about 50 billion parameters per token. The practical result: 300+ output tokens per second, 5x higher throughput than previous versions, and roughly 30% lower inference costs than competing models at similar capability levels. NVIDIA built it for “advanced reasoning, planning, and agentic workflows,” which is corporate-speak for “this model can plan multi-step tasks and execute them without hand-holding.”

Why does a hardware company releasing an AI model matter for founders? Pricing pressure. Meta’s Llama models already forced the open-source AI conversation. NVIDIA entering with a model trained on its own hardware and optimized for its own inference stack adds another option that’s specifically designed to run cheaply on NVIDIA infrastructure. If you’re choosing between paying OpenAI’s API pricing, running Meta’s Llama on rented GPUs, or running Nemotron 3 on NVIDIA’s stack, you now have a third path that’s vertically integrated from chip to model.

The model will be available on Hugging Face, ModelScope, OpenRouter, and as an NVIDIA NIM microservice. For founders already using AI tools for their businesses, this means more competition among model providers, which means lower prices and more options. That’s pure upside regardless of which model you end up using.

The Super tier of the Nemotron 3 family, launched in March 2026 at 120 billion parameters, targets mid-range enterprise applications. Notice that number: 120 billion parameters. The same size RTX Spark can run locally. That’s not a coincidence. NVIDIA designed a model family where the mid-tier runs on its laptop chip and the top tier runs on its data center hardware. Full-stack infrastructure means the software and hardware are designed together.

Three things founders should do right now

The announcements are concrete enough to act on. Here’s what makes sense today, not in some theoretical future.

Audit your AI inference costs. Pull your last three months of API spend from OpenAI, Anthropic, or whatever provider you use. Calculate what percentage of your workloads use models at or below 120 billion parameters. That’s your RTX Spark-eligible spend, the portion you could potentially move to local inference when these laptops ship in fall 2026. If it’s more than $300 per month, the hardware pays for itself within a year.

Test your code on CUDA. RTX Spark runs the same CUDA binaries as H100 data center hardware. If your AI pipeline already uses PyTorch with CUDA, you’re set. If it doesn’t, if you’ve been running on CPU-only or using a non-NVIDIA inference stack, now is the time to evaluate whether porting to CUDA makes sense. The CUDA advantages are compounding: same code runs on laptop, workstation, and data center.

Move your agent roadmap forward. Every signal from Computex pointed the same direction: agentic AI has production-grade hardware underneath it now. Vera CPU for server-side agents. RTX Spark for client-side agents. Nemotron 3 for the models. Cosmos 3 for physical world agents. If your product includes AI agents planned for 2027, evaluate whether the infrastructure timeline justifies building them sooner. GitHub activity already shows the shift: code commits roughly tripled in early 2026 compared to 2025 levels, driven largely by AI-assisted and AI-agent development.

The founders who’ll benefit most from NVIDIA’s infrastructure pivot are the ones who recognize what it actually is: a bet that AI computing demand will grow exponentially, and the company that provides the plumbing profits from every gallon that flows through. Whether you’re building solo AI businesses or scaling a team, understanding who controls your infrastructure layer is now as important as choosing your tech stack.

Read More From the GROW desk