📚 Explainer
AISemiconductor

What is Physical AI — and Why Does the Next Decade of the AI Revolution Happen in the Physical World?

Every previous wave of artificial intelligence has lived behind a screen. ChatGPT processes language. Stable Diffusion generates images. AlphaFold predicts proteins. None of them can pick up a box, weld a joint, or navigate a factory floor without human direction. Physical AI is the term for what comes next — AI systems that perceive, reason about, and act in the physical world. The robots being built right now in Shanghai, Austin, and Tokyo are the first infrastructure of that transition.

What physical AI is — and what it requires. Physical AI, in Jensen Huang's definition from CES 2026, is "AI systems that understand physical laws and interact with the physical world." The concept has older roots — robotics researchers have worked on embodied intelligence for decades — but what has changed in 2025–26 is the convergence of three capabilities that previously did not exist at deployable quality simultaneously: models that can perceive and reason in three dimensions, edge compute powerful enough to run them in real time, and simulation environments realistic enough to generate the training data that generalises from virtual environments to real ones.

Physical AI requires things that screen-based AI does not. Perception: cameras, LiDAR, tactile sensors, and inertial measurement units that map the environment and the robot's own position within it in real time. Action: actuators, motors, and joints that translate AI decisions into physical movement with enough precision and force to be useful in an industrial setting. And the bridge: foundation models trained on vast physical interaction data that can generalise from what they learned in simulation to what they encounter in uncontrolled deployment. That last element — closing the sim-to-real gap — remains the field's defining challenge.

The model breakthrough that made it real. The dominant technical approach in 2026 uses Vision-Language-Action (VLA) models — neural networks that take camera input and language instructions and output robot actions directly. ICLR 2026 received 164 VLA paper submissions, an 18× increase from one year prior. NVIDIA's GR00T N1.6, announced at CES 2026, is a 32-layer diffusion transformer trained on thousands of hours of teleoperation data across multiple robot bodies. Physical Intelligence's pi‑0.5 demonstrated meaningful open-world generalisation across 68 tasks on seven different robot platforms. These are not incremental improvements to industrial controllers — they are the same foundation model architecture that produced ChatGPT, adapted to perceive and act rather than predict text.

NVIDIA's Cosmos 3.0, previewed at GTC 2026, is the first world foundation model that unifies synthetic world generation, vision reasoning, and action simulation in a single architecture. It generates physically plausible synthetic training data at scale, directly addressing the data scarcity that has constrained robot learning since the beginning. NVIDIA's CEO declared at GTC 2026 that "every industrial company will become a robotics company" and announced that the big bang of physical AI had arrived, with $20 billion invested in humanoid robots to that point. The comparison to the arrival of ChatGPT in 2022 was not rhetorical.

EXHIBIT 1

The physical AI market — already larger than most people realise, growing faster than consensus expects

Total physical AI market size including industrial automation, autonomous vehicles, and humanoid robots, $ billion; 2026F onward are forecasts

$0 $100B $200B $300B $400B $35B 2023 $55B 2024 $81B 2025 (confirmed) $108B 2026F $192B 2028F $430B 2030F ~33% CAGR 2025 → 2030F Jensen Huang (NVIDIA): “humanoid robots are a $40 trillion TAM” — labour automation long-run Confirmed data Forecast (33% CAGR)

Source: Kaiso Research (Physical AI Market 2026–2035, $81.4B confirmed 2025, 33.49% CAGR); FutureMarkets (physical AI surpasses $430B by 2030); Goldman Sachs ($50B humanoid robotics investment by 2030); ATF extrapolation for 2026F and 2028F. Market definition: industrial automation + autonomous vehicles + humanoid robots + healthcare robotics. Embodied AI sub-market (MarketsandMarkets): $4.44B (2025) → $23.06B (2030), 39% CAGR.

NVIDIA: positioning as the CUDA of robots. NVIDIA's strategy for physical AI is structurally identical to its strategy for generative AI: own the training compute, own the simulation environment, own the inference processor, and collect rent on every model trained and every robot deployed. At GTC 2026, NVIDIA announced Cosmos 3.0 — its world foundation model — alongside partnerships with ABB Robotics, FANUC, AGIBOT, Agility Robotics, and LG Electronics. Its reference humanoid "Isaac Root" stands 6 feet, weighs 150 pounds, has 31 degrees of freedom and 25 per hand, and is built in partnership with Unitree. The Jetson T4000, announced at CES 2026 at $1,999 per unit, delivers 1,200 TFLOPS within a 40–70W power envelope — the edge compute that makes real-time robot inference economically viable at manufacturing scale.

Tesla and China racing for production volume. Tesla's Optimus Gen 3 entered mass production at Fremont in January 2026, with $20 billion in capital expenditure committed to humanoid output this year. The 2026 target is 100,000 units — a number that, if achieved, would dwarf all other manufacturers combined. China is not waiting. AgiBot has shipped 5,100 humanoid units with 39% global market share, operates a 3,000-square-metre Giga Data Factory in Shanghai where hundreds of robots are teleoperated to generate training data, and has released Lingqu OS — an embodied intelligence operating system. BYD targets 20,000 humanoid deployments in its own EV and battery manufacturing lines. TrendForce projects China will ship approximately 62,500 humanoid units in 2026, a 94% increase year-on-year, representing over 80% of global production.

EXHIBIT 2

The physical AI technology stack — and who owns each layer

Five-layer architecture from physical sensing to robot deployment; upward arrows show the training data and model flow

1  ·  PHYSICAL WORLD & SENSORS LiDAR  ·  RGB cameras  ·  tactile sensors  ·  IMU  ·  force-torque sensors Robot hardware supply Japan (actuators) · China (arms) 2  ·  DATA COLLECTION Teleoperation  ·  AgiBot AIDEA Giga Factory  ·  sim-to-real augmentation  ·  OSMO workflow NVIDIA Isaac ROS · OSMO 3  ·  SIMULATION & WORLD MODELS NVIDIA Cosmos 3.0  ·  Omniverse  ·  Isaac Lab-Arena  ·  physics-based synthetic data generation NVIDIA Cosmos 3.0 · Omniverse 4  ·  FOUNDATION MODELS (VLA) GR00T N1.6 (NVIDIA)  ·  pi-0.5 (Physical Intelligence)  ·  GR00T N2 (preview)  ·  Cosmos Predict 2.5 NVIDIA GR00T · Isaac GR00T N 5  ·  EDGE INFERENCE & DEPLOYMENT Jetson T4000 (1,200 TFLOPS · $1,999)  ·  Jetson Thor  ·  Robot body  ·  Real-world operation NVIDIA Jetson T4000 · Thor NVIDIA FULL STACK

Source: NVIDIA GTC 2026 keynote (Cosmos 3.0, GR00T N2 preview, Isaac partnerships); NVIDIA CES 2026 (GR00T N1.6, Jetson T4000); AgiBot AIDEA Giga Factory (Reuters, Shanghai); Physical Intelligence pi-0.5 (company release); NVIDIA Cosmos Predict 2.5 / Transfer 2.5 (CES Jan 2026). VLA = Vision-Language-Action model, the dominant paradigm for robot intelligence in 2026.

“The ChatGPT moment for robotics is here. Breakthroughs in physical AI — models that understand the real world, reason and plan actions — are unlocking entirely new applications.”

Jensen Huang, Founder and CEO, NVIDIA — CES 2026 / GTC 2026

Asia's structural advantages in the physical AI race. The hardware layer of physical AI plays directly to Asia's manufacturing strengths in a way that software-only AI never did. Harmonic drives — the precision gear reducers that give robot joints their accuracy and load capacity — are dominated by Japan's Harmonic Drive Systems and Sumitomo Heavy Industries. Brushless motors, torque sensors, and the actuator components that define a robot's physical capability are concentrated in Japan and increasingly in Shenzhen, where Unitree produces sub-$30,000 humanoids using a supply chain density that cannot be replicated outside the Pearl River Delta.

Japan has committed $65 billion to capture over 30% of the global robotics market by 2040, recognising that physical AI deployment is a structural match for an economy with an ageing workforce, world-leading factory automation heritage, and deep semiconductor and materials supply chains. BYD, CATL, and Foxconn — the world's largest demand pools for manufacturing labour — are not just customers of physical AI. They are the case study that will determine whether it works at industrial scale.

Three predictions for 2026–27. First: by end-2027, cumulative humanoid deployments in Chinese EV and battery factories exceed 100,000 unit-instances, making China's auto sector the first industry to operate humanoids at genuine commercial scale. BYD and CATL's closed, known-environment production lines are structurally better suited to early deployment than the uncontrolled environments where general-purpose robots struggle.

Second: NVIDIA's Cosmos simulation platform becomes for physical AI what CUDA became for deep learning — the default training environment for 70%+ of robot development workflows by 2028. The data generation flywheel (more simulation data → better models → more deployments → more real-world data → better simulation) is already running, and NVIDIA controls the simulation environment that feeds it.

Third: a major Japanese robotics integrator — FANUC, Yaskawa, or a corporate spinout — announces a deep strategic alliance with a foundation model provider by mid-2027, accelerating the integration of VLA models into Japan's installed base of 276,000 annual industrial robot installations.

EXHIBIT 3

The 2026 humanoid production race — volume targets from leading manufacturers

2026 humanoid unit production targets or projections; scale set at 30k to show relative ambitions — Tesla bar extends beyond scale

0 10,000 20,000 30,000+ Humanoid units (2026 target / projection) Tesla Optimus Gen 3 (USA) — $20B capex in 2026 100,000 (target) China (all manufacturers combined) — TrendForce ≈62,500 BYD (China) — EV & battery factory deployment 20,000 (target) AgiBot (China) — 39% global market share ~10,000 (target) USA 🇺🇸 China 🇨🇳 China 🇨🇳 China 🇨🇳 Tesla and China total bars extend beyond 30k scale. All figures are targets or projections; actual 2026 shipments may vary materially.

Source: Tesla $20B capex commitment and 100,000-unit 2026 target (Reuters / Tesla IR); TrendForce China humanoid robot shipments projection ~62,500 units (+94% YoY, 2026); BYD humanoid target (Futuremarkets / industry disclosures); AgiBot 5,100 units shipped + 39% market share (RaisesSummit, RCSV Research China Robotics 2026). Targets represent stated company goals, not confirmed production outcomes. China total includes Unitree, UBTECH, Fourier, Agibot, and others.

The Bottom Line

Physical AI is not a distant promise — it is a production ramp. BYD factories in Shenzhen are deploying humanoids on battery assembly lines. Tesla's Fremont plant has converted automotive production lines to humanoid manufacturing. NVIDIA is building the CUDA of robots. The next decade of the AI trade is not about making language models bigger. It is about whether machines can learn to handle the physical world reliably enough to replace structured human labour at scale. That question is being answered right now, in factories across Asia — and the answer, for specific environments and specific tasks, is increasingly yes.

Colin Tan  ·  Editor, Asia Tech Feed

Colin covers semiconductors, AI infrastructure and supply-chain dynamics across the Asia-Pacific region. He has tracked the embodied AI and humanoid robotics industry since NVIDIA's first GR00T announcement and writes the daily ATF digest. Reach him at [email protected] or connect on LinkedIn.