The Brief
NVIDIA CEO Jensen Huang opened GTC 2026 in San Jose by unveiling the full Vera Rubin platform — six new chips headlined by the R100 GPU, which pairs 336 billion transistors with 288 GB of HBM4 memory and delivers up to 50 petaflops of FP4 inference per unit, roughly five times Blackwell’s output. Mass production is expected in the second half of 2026, with an Ultra variant and 15-exaflop NVL576 racks following in 2027.
The Report
Jensen Huang took the stage at San Jose’s SAP Center on Monday to formally launch the Vera Rubin architecture, NVIDIA’s successor to Blackwell and the company’s most aggressive generational leap in AI compute to date. The platform encompasses six purpose-built chips: the Rubin GPU, the Vera CPU, an NVLink 6 switch, the ConnectX-9 SuperNIC, the BlueField-4 DPU, and a Spectrum-6 Ethernet switch.
The centrepiece R100 GPU is built on TSMC’s 3nm process using a dual-die design — two reticle-sized compute chiplets containing 336 billion transistors. Each GPU carries 288 GB of HBM4 memory at 22 TB/s bandwidth, nearly tripling Blackwell’s 8 TB/s. NVIDIA claims 50 petaflops of FP4 inference and 35 petaflops of FP4 training per GPU, representing a 5.6x and 3.5x improvement over the B200 respectively. The accompanying Vera CPU uses 88 custom Arm-based Olympus cores with 176 threads, up to 1.5 TB of LPDDR5x memory, and is the first NVIDIA CPU to support FP8 precision — a design choice that reflects the growing computational demands of agentic AI workloads. “CPU performance has become a key bottleneck in scaling AI agent tasks,” NVIDIA’s head of AI infrastructure Dion Harris noted.
At the rack level, the NVL72 system — 72 GPUs and 36 CPUs — delivers 3.6 exaflops of FP4 inference with 20.7 TB of total HBM4 and 260 TB/s of NVLink bandwidth, a 3.3x performance increase over the current GB300 NVL72. A larger NVL144 CPX configuration, pairing standard Rubin GPUs with the new Rubin CPX prefill-phase inference accelerator, reaches 8 exaflops. NVIDIA projects the CPX platform will generate $5 billion in token revenue for every $100 million in capital expenditure.
The economic claims are equally pointed: a 10x reduction in inference token cost versus Blackwell, four times fewer GPUs needed to train mixture-of-experts models, and 18 times faster assembly and servicing. Independent benchmarks from MLPerf are not yet available; all performance figures are NVIDIA’s own published specifications.
Beyond silicon, Huang previewed NemoClaw, an open-source enterprise AI agent platform designed to orchestrate models into task-executing agents across internal systems. The platform is hardware-agnostic, though its deep integration with NVIDIA’s software stack creates a familiar gravitational pull. NVIDIA has begun pitching partnerships with Salesforce, Cisco, Google, Adobe, and CrowdStrike.
The conference also carried forward NVIDIA’s $20 billion Groq licensing agreement from December, which brought inference-specialist personnel and architecture into the fold, and a multiyear gigawatt-scale partnership with Thinking Machines Lab. GTC itself has drawn 30,000 attendees from 190 countries across more than 700 sessions, with physical AI, quantum computing, and robotics featuring prominently alongside the core infrastructure announcements.
Huang framed the moment in characteristically expansive terms: “Every company will use it. Every nation will build it.” NVIDIA reported fiscal year 2026 revenue of $215.9 billion, a 65% year-over-year increase, with data centre operations accounting for more than 85% of total revenue. The company’s market capitalisation stands at approximately $4.5 trillion. Rubin-based products are expected in the second half of 2026, with the Vera Rubin Ultra — a four-die-per-GPU variant delivering 100 petaflops per package and 15 exaflops per NVL576 rack — slated for late 2027.
The Angle
The numbers are large enough to produce a kind of anaesthesia. Exaflops, petabytes per second, 336 billion transistors — at a certain density, specifications stop functioning as information and start functioning as atmosphere. Which is useful if you are NVIDIA, because the atmosphere is the product. What Huang is actually selling in San Jose is not a chip. It is the assumption that the trajectory is fixed, the demand is permanent, and the only variable is who builds fast enough to stay on the curve.
That assumption deserves more scrutiny than it is receiving. Not because it is wrong — the evidence for sustained inference demand is strong and getting stronger — but because of what it makes invisible. The FP64 performance on Rubin dropped from Blackwell’s 45 teraflops to 33. That regression is absent from every keynote slide and most coverage. It is a small number in the wrong direction, and it reveals the actual strategic bet: NVIDIA has decided that the future of compute is overwhelmingly inference at low precision, not high-performance scientific simulation at double precision. The architecture has been optimised for generating tokens, not modelling physics. For a company that began as a graphics processor manufacturer and grew through HPC, this is not a minor rebalancing. It is a declaration about what computing is for now.
The NemoClaw announcement carries a similar quiet signal beneath the open-source branding. A hardware company building an application-layer agent platform and giving it away free is not philanthropy. It is the recognition that the margin is moving — from selling picks to owning the path between the mine and the market. Every enterprise agent workflow that runs on NemoClaw is a workflow optimised for NVIDIA silicon by default, whether or not it is required to be. The hardware-agnostic label is technically accurate and strategically irrelevant.
AMD’s Helios offers 50% more memory per GPU and comparable exaflops. Intel’s Gaudi 4 exists. Hyperscaler custom silicon — Amazon’s Trainium, Google’s TPUs — continues to scale. None of it has dented NVIDIA’s 85% market share, because the competition is not selling chips into a market. It is selling chips into an ecosystem six million developers deep, with a software stack that has been compounding for fifteen years. The moat is not performance. The moat is CUDA, and everything Huang announced today — the hardware, the agents, the partnerships — is designed to deepen it before anyone finds a way across.
The $5 billion in projected token revenue per $100 million in capital expenditure is the number that will matter most when this conference is remembered. It is NVIDIA pricing the future of intelligence as a utility — metered, margined, and mediated through its infrastructure. Whether that pricing holds depends on whether inference remains as GPU-bound as training was. The Groq acquisition suggests NVIDIA is not entirely certain it will.
The company that built the tools is now building the factory, the supply chain, and the storefront. The question is no longer who makes the best chip. It is who owns the infrastructure through which intelligence gets delivered — and what it costs everyone else to rent it.