The Brief

Nvidia’s GPU Technology Conference opens Monday in San Jose with over 30,000 attendees expected across ten venues, as CEO Jensen Huang prepares a keynote widely anticipated to debut a dedicated inference processor built on technology acquired from Groq and an open-source enterprise AI agent platform called NemoClaw. The four-day event arrives as the AI industry’s centre of gravity shifts decisively from model training to deployment at scale — a transition that threatens to redraw the competitive map Nvidia currently dominates.

The Report

Jensen Huang takes the stage at the SAP Center on Monday at 11 a.m. Pacific Time for what has become the AI industry’s most consequential annual address. GTC 2026 runs March 16–19, spanning ten venues across downtown San Jose, with over a thousand sessions, seventy hands-on labs, and full conference passes already sold out.

The headline expectation is a new inference-optimised chip incorporating technology from Groq, the inference startup whose deterministic, SRAM-based architecture Nvidia licensed for $20 billion in December 2025. The deal — structured as a non-exclusive licence and acqui-hire to avoid triggering mandatory merger reviews — brought Groq founder Jonathan Ross and president Sunny Madra into Nvidia alongside much of the engineering team. According to the Wall Street Journal, the resulting chip integrates Groq’s Language Processing Unit architecture with Nvidia’s fabrication capabilities on TSMC’s A16 process node with 3D stacking. OpenAI has received early access and committed 3 gigawatts of dedicated inference capacity, with the chip initially targeting its Codex programming tool.

The second major expected reveal is NemoClaw, an open-source enterprise AI agent platform that allows companies to deploy autonomous AI agents regardless of whether their infrastructure runs on Nvidia hardware. Nvidia has reportedly pitched the platform to Salesforce, Cisco, and Google, positioning it as a security-hardened alternative to OpenClaw after documented incidents — including an AI agent mass-deleting emails during Meta research — raised enterprise concerns about agent reliability.

The announcements sit within the broader architecture of the Vera Rubin platform, unveiled at CES in January: six integrated chips delivering a claimed 10x reduction in inference token cost over Blackwell and 50 petaflops of inference compute per GPU. Cloud deployments through AWS, Google Cloud, Microsoft, and others are scheduled for the second half of 2026. Looking further ahead, Huang has teased “several new chips the world has never seen before,” with analysts speculating about a silicon photonics breakthrough and the N1X AI PC superchip — a joint venture with MediaTek targeting the high-end laptop market.

The competitive backdrop lends the keynote additional weight. Inference now accounts for approximately two-thirds of all AI compute demand, up from roughly one-third in 2023. Google, Amazon, Cerebras, and AMD have all expanded their inference-specific offerings. Specialised inference ASICs are projected to capture 45 percent of the inference market by 2030. Nvidia commands an estimated 80 to 95 percent of the training chip market, but training is no longer where the growth is.

Wall Street remains broadly confident. Ninety-three percent of covering analysts rate the stock a buy, with price targets ranging from $267 to $360 against a current price near $185. Bank of America’s Vivek Arya expects Nvidia to provide three-generation product visibility stretching to 2028’s Feynman architecture — a roadmap depth he argues “locks in developer and enterprise commitments well ahead of rivals.”

Huang’s keynote livestreams free at nvidia.com. An analyst Q&A follows Tuesday morning.


The Angle

The interesting thing about GTC 2026 is not what Nvidia is announcing. It is what the announcements admit. A company that spent the past three years as the undisputed supplier of the hardware that trains the world’s most capable AI models is now spending $20 billion and an entire product cycle to make sure it also owns the hardware that runs them. That is not confidence. That is a company watching the market move beneath it and recalculating in real time.

The Groq deal is the tell. Nvidia’s GPU architecture was built for training — brute-force matrix multiplication at enormous scale, where latency matters less than throughput. Inference is structurally different: thousands of lightweight calls per second, each needing to return fast enough that a user or an autonomous agent doesn’t stall. Groq’s SRAM-based architecture was purpose-built for exactly that workload, achieving internal memory bandwidth an order of magnitude beyond what HBM can deliver. Nvidia did not license this technology because it was interesting. It licensed it because the alternative was ceding the fastest-growing segment of AI compute to companies that had already built the right tool for the job.

NemoClaw points in the same direction, from the software side. Making the agent platform hardware-agnostic is an unusual move for a company whose competitive moat has always been ecosystem lock-in through CUDA. The calculation is legible: if AI agents become the primary consumer of inference compute — and the trajectory suggests they will — then the company that controls the orchestration layer captures value regardless of whose silicon sits underneath. It is a hedge dressed as generosity.

None of this diminishes Nvidia’s position. The Vera Rubin platform is a genuine engineering achievement, and the one-year architecture cadence from Blackwell to Rubin to Feynman is a pace no competitor has matched. But the shape of the strategy has shifted. Two years ago, Nvidia was selling the picks during a gold rush. Now it is trying to own the mine, the assay office, and the transport routes simultaneously — because it has realised the gold rush is moving to a different mountain, and the picks it already sells are not the right ones for the new terrain.


The company that built the infrastructure for teaching machines to think is now racing to build the infrastructure for letting them act. The difference between those two problems is the difference between this decade and the next.