NVIDIA founder and CEO Jensen Huang: AI factories, physical AI and the robotics revolution

CN
4 hours ago

Written by: Techub News Compilation

Introduction

In January 2025, NVIDIA founder and CEO Jensen Huang delivered a keynote speech lasting 91 minutes at CES (Consumer Electronics Show). This speech not only serves as a barometer for the consumer electronics sector but also as a significant proclamation for the AI and computing industry. A year after the generative AI wave swept the globe, Huang systematically elaborated on NVIDIA's full-stack strategy from chips, software to ecosystem, and unveiled the grand blueprint for AI moving from the digital world into the physical world. The backdrop of this speech is the full production of Blackwell architecture GPUs and the penetration of AI from the cloud to the edge, PCs, and even the core of every robot.

Summary

  • Release of Blackwell architecture GeForce RTX 50 series GPUs, with significant performance improvements, making AI-driven neural rendering the core of future computer graphics.
  • Unprecedented scaling of AI factories, with Blackwell systems in production across 45 factories globally to meet the massive computational demands driven by three major expansion laws.
  • Launch of the world's first physical AI foundational model platform, Cosmos, aimed at understanding the physical world and providing a "world model" for robotics, autonomous driving, and more.
  • Prediction that intelligent agent AI, autonomous vehicles, and humanoid robots will form the future multi-trillion dollar robotics industry, along with announcements of corresponding technology stacks and partner progress.
  • Introduction of Project Digits, a new type of desktop AI supercomputer designed to make powerful AI computing accessible to every developer.

From GeForce to Blackwell: How AI Reshapes Graphics and Gaming

At the beginning of the speech, Huang reflected on NVIDIA's 30-year journey, from the initial gaming graphics processors (GPUs) to the engine that drives the AI revolution today. He specifically pointed out that it was the GeForce platform that brought AI to the masses, and now AI is "giving back" and thoroughly revolutionizing GeForce itself.

A segment of real-time rendering showcasing stunningly detailed game visuals was demonstrated on-site. Huang explained that behind this is the AI-driven neural rendering technology—the latest leap in DLSS (Deep Learning Super Sampling). Traditional ray tracing requires computing each pixel, resulting in massive computational loads. The new generation of DLSS combines programmable shaders, ray tracing acceleration, and AI prediction. The system renders only a few key pixels, and then a neural network trained on supercomputers predicts and generates all the remaining pixels based on these, even predicting future frames. Huang emphasized its revolutionary nature with a set of numbers: at 4K resolution, each frame contains about 33 million pixels, but the system only computes 2 million of them, with the rest efficiently generated by AI. "This is truly a miracle," he said, "The computational burden for AI is much lower, but after massive training, the generation efficiency is extremely high." This marks the future of computer graphics as neural rendering, the fusion of artificial intelligence and computer graphics.

On this basis, Huang officially announced the release of the GeForce RTX 50 series GPUs based on the Blackwell architecture. The flagship model RTX 5090 has 92 billion transistors and offers 4 PetaFLOPS of AI performance (three times that of the previous generation Ada) and 380 TeraFLOPS of ray tracing performance. It uses Micron's G7 memory with a bandwidth of 1.8 TB/s. Huang particularly mentioned that the shaders of this generation of GPUs can also process neural networks, enabling "neural texture compression" and "neural material shading," resulting in unprecedented image quality. The RTX 5090 is priced at $1599, which he humorously referred to as "a great investment for your $10,000 home entertainment command center."

Even more impressive is that such powerful performance has been packed into laptops. Huang showcased the RTX 570 laptop, whose performance rivals the previous generation desktop flagship RTX 4090, starting at just $1299. This is made possible by AI neural rendering significantly enhancing energy efficiency, allowing high-performance GPUs to fit into slim notebooks. He concluded, "GeForce has brought AI to the world and democratized it; now, AI has come back to thoroughly revolutionize GeForce."

AI Factories and the Three Major Expansion Laws: The Perpetual Engine Driving Demand for Computing Power

Subsequently, Huang shifted the topic to the broader AI industry. He pointed out that the entire industry is chasing and competing to expand the scale of artificial intelligence, driven by strong "expansion laws."

He elaborated on the three major expansion laws currently observed:

  • Pre-training Expansion Law: Model capabilities improve with increased volumes of training data, model scale, and computational resources used. The multimodal data generated by the Internet (videos, images, sounds) is still growing rapidly, providing fuel for training foundational models.
  • Post-training Expansion Law: Through techniques such as reinforcement learning and human feedback, AI continues to refine its skills after "graduation," such as solving mathematical problems and optimizing reasoning abilities. This is similar to coaching for targeted tutoring.
  • Testing Time Expansion Law: When AI is in use, it can dynamically allocate computational resources to solve problems. For example, by using "long thinking" or "chain of thought" to break complex problems into multi-step reasoning, generating multiple plans and evaluating the optimal solution. This greatly improves answer quality but also significantly increases the computational demands during the reasoning phase.

Huang emphasized that these laws collectively drive a "huge demand" for NVIDIA's computing platforms, particularly for Blackwell chips. He announced that Blackwell is now in full production. All major global cloud service providers have deployed Blackwell-based systems, with about 15 computer manufacturers producing more than 200 different configurations of servers (air-cooled, liquid-cooled, x86, or NVIDIA Grace CPU versions).

To visually demonstrate the scale of AI factories, Huang had staff bring up a complete GB200 NVL72 system. This machine weighs 1.2 tons, contains approximately 600,000 parts (equivalent to 20 cars), consumes 120 kilowatts of power, and internally has two miles of copper cables and 5,000 wires. It is being manufactured, tested, and assembled in 45 factories worldwide. "The manufacturing process is astonishing," he said, "but the goal of all this is that the expansion laws are driving computational demand so powerfully."

Compared to the previous generation Hopper, Blackwell has improved energy efficiency by 4 times and cost-effectiveness by 3 times. This means the cost of training models has been reduced by two-thirds, or in other words, one can train a model three times larger for the same cost. Huang pointed out that every data center is constrained by electrical power, so a fourfold improvement in energy efficiency means that the AI "tokens" generated by the data center and the business revenue it can support also significantly increase accordingly. "These AI factory systems are truly the factories of today."

Intelligent Agent AI: The Next Generation of Digital Workforce and Corporate Software Revolution

Huang predicted that intelligent agent AI will be the next robotics-level industry, potentially creating trillions of dollars in value. Intelligent agent AI is not a single model but a system composed of multiple models that can perceive, reason, plan, and act to perform tasks on behalf of humans.

To help build the ecosystem for intelligent agent AI, NVIDIA has launched a three-layer technology stack:

  1. NVIDIA NIM: AI microservices. This packages complex CUDA software and optimized models into containers that developers can easily deploy on any cloud or local system. NIM provides models for vision, language, speech, animation, digital biology, and more.
  2. NVIDIA Nemo: Digital employee onboarding and training framework. In the future, AI intelligent agents will become a "digital workforce" for enterprises, and Nemo provides tools to customize agents (learning company-specific terminology and business processes), assess their performance, set behavior guardrails, and manage their lifecycle. Huang predicted, "In the future, every company's IT department will become the human resources department for AI intelligent agents."
  3. Blueprints and Open Source Models: NVIDIA offers various reference blueprints for intelligent agent applications and announces the launch of the "Llama Neotron" series of open-source models based on fine-tuning of Meta Llama 3.1. These models are optimized for enterprise use cases and excel in multiple benchmark tests, becoming a significant cornerstone for building enterprise-grade intelligent agents.

He listed multiple application scenarios for intelligent agent AI: providing AI research assistants for knowledge workers, automatically analyzing complex documents and generating podcasts; helping developers continuously scan for code security vulnerabilities; assisting researchers in quickly screening billions of compounds to discover candidate drugs; analyzing massive video data through city camera networks for traffic monitoring, event alerts, and automated reporting. Huang summarized, "The era of intelligent agent AI has arrived, suitable for every organization."

Physical AI and Cosmos: Endowing Robots with a "World Model"

One of the most forward-looking parts of the speech was about "physical AI." Huang pointed out that current large language models deal with text tokens, but if the input is a physical environment (like sensor data), the request is a physical action (like "go pick up that box"), and the output is action tokens, then we enter the realm of robotics.

Robots need a "foundational model" that understands the physical world. It needs to grasp gravity, friction, inertia, geometric spatial relationships, causality, and object permanence, among others. To this end, NVIDIA has launched the world's first physical AI foundational model platform—Cosmos.

Cosmos is trained using 20 million hours of videos focused on physical dynamics (such as natural phenomena, human movements, and object manipulation), aiming to teach AI to understand the physical world. It includes autoregressive models (for real-time applications), diffusion models (for high-quality image generation), advanced tokenizers, and an end-to-end data processing pipeline accelerated by CUDA and AI.

Huang emphasized the importance of the integration of Cosmos with NVIDIA Omniverse (a physics-based digital twin simulation platform). Omniverse provides physical "foundational facts" that can control and constrain the generation of Cosmos, ensuring that its outputs comply with physical laws, similar to using retrieval-augmented generation (RAG) to ground large language models in real information. The combination creates a "physics-based multiverse simulator" that can be used for synthetic data generation, model training, and testing.

He proposed a "three computer" architecture that future robotic companies must have:

  • DGX: For training AI models.
  • Omniverse + Cosmos: A digital twin simulation platform for testing, validating AI, and generating massive synthetic training data.
  • AGX (like Thor): Onboard supercomputers deployed on edge devices such as robots and vehicles.

The Robotics Revolution: Autonomous Driving, Humanoid Robots, and Industrial Digitalization

Based on the "physical AI" and "three computer" framework, Huang painted a picture of the upcoming robotics revolution, focusing on three key areas: intelligent agent AI (information workers), autonomous vehicles, and humanoid robots. He predicted that these three will constitute "the largest technological industry ever in history."

1. Autonomous Vehicles: Huang announced that NVIDIA's next-generation onboard computer Thor has started full production. Thor's processing power is 20 times that of the previous generation Orin, and its operating system Drive OS has received the highest level of automotive functional safety certification (ASIL D). He revealed that NVIDIA is collaborating with nearly all mainstream automotive companies and specifically announced a partnership with Toyota in the field of next-generation autonomous driving.

He demonstrated how to use Omniverse and Cosmos to build an "autonomous vehicle data factory": reconstructing real road data in a digital twin environment, and then using AI to generate nearly infinite, physics-compliant synthetic driving scenarios (such as varying weather, lighting, and traffic conditions), thereby expanding thousands of miles of real road test data into billions of miles of effective training data. "We will have an abundance of training data," Huang said, "the autonomous driving industry has arrived."

2. Industrial Digitalization and Warehousing Logistics: NVIDIA is working with global leaders in warehouse automation, Körber, and Accenture to bring physical AI to the trillion-dollar scale warehousing and logistics market using the Omniverse blueprint "Mega." By building a digital twin of the warehouse, companies can simulate and optimize robot fleet scheduling, task allocation, and predict key operational metrics before making physical transformations.

3. Humanoid Robots: Huang believes that the "ChatGPT moment" for general-purpose robots is approaching. He introduced the NVIDIA Isaac Groot robot development platform, which provides foundational models for robots, data pipelines, simulation frameworks, and Thor robot computers. Through the "synthetic motion generation" workflow, developers can conduct a small amount of manual remote operation demonstrations using devices like Apple Vision Pro, and then expand these demonstrations into massive synthetic training data in Omniverse and Cosmos, efficiently training robot strategies. "We will have an abundance of data to train robots," he summarized.

Project Digits: Bringing AI Supercomputers to Every Desktop

At the end of the speech, Huang returned to NVIDIA's starting point: making powerful computing capabilities more accessible. He reminisced about the delivery of the first DGX-1 AI supercomputer to OpenAI in 2016, which ushered in a new era of AI research. Today, AI has become mainstream, and every developer and creator needs an AI supercomputer.

Thus, he unveiled "Project Digits"—a new type of desktop AI supercomputer. This device is compact in appearance and includes a secret chip "GB110" developed in cooperation with MediaTek (a miniaturized Grace Blackwell chip), with CPU and GPU connected via NVLink between the chips. It runs a complete NVIDIA AI software stack, serving as both a local workstation and remote access like a cloud supercomputer. Huang stated that Project Digits is expected to be released around May 2025. "Imagine, who wouldn't want a device like this?" he laughed.

Finally, Huang concluded the speech with a video of reflection and outlook, thanking partners and wishing everyone a Happy New Year. The entire speech was not just a product launch but a systematic declaration on how AI is transitioning from digital computing to the physical world, reshaping trillion-dollar industries.

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink