Perspective

Market Deep Dive: Physical AI

By Nymeria
TL;DR
  • Core Thesis: Physical AI is a capital-intensive, multi-layer ecosystem where investment is concentrating in the intelligence layer (foundation models) far more than in hardware or enabling infrastructure, creating structural asymmetries in undercapitalized segments.
  • Why It Matters: Disclosed Physical AI equity funding reached USD 8.73 billion, with the median round at USD 112.5 million. This is not traditional robotics pacing. It is frontier AI infrastructure pacing applied to the physical world.
  • Strategic Direction: Capital is overwhelmingly concentrated in cross-embodiment intelligence platforms. Enabling layers such as simulation infrastructure, developer tooling, and data pipelines are systematically underfunded relative to their structural importance to the Physical AI stack.

When Jensen Huang took the GTC stage in 2026, he delivered a declarative statement rather than a prediction. "Physical AI has arrived," he said. "Every industrial company will become a robotics company." The numbers beneath that claim reflect a structural shift. Robotics and Physical AI attracted approximately USD 40.7 billion in global investment in 2025, outpacing every other deep-tech category tracked by CB Insights.

But Physical AI is not a single market. It is an umbrella term spanning robotic foundation models, vision-language-action architectures, world models, humanoid hardware, general-purpose robot platforms, simulation environments, synthetic data pipelines, and developer tooling for autonomous systems. Each layer of this stack has different capital requirements, different technology maturity, and a different competitive timeline. Mapping this landscape accurately requires separating these layers by function and funding concentration. We have previously analysed one critical enabling layer, the simulation, digital twin, and world model infrastructure for industrial operations, in our Market Deep Dive: Industrial Simulation. This analysis takes the full stack view.


The Physical AI Landscape

The central challenge facing anyone mapping Physical AI is a category problem rather than a technology problem. Physical AI has become an umbrella label applied across robot foundation model labs, humanoid hardware startups, industrial simulation platforms, sensor manufacturers, and autonomous system builders. Grouping them under a single market label obscures more than it reveals. The intelligence layer, the embodiment layer, and the enabling infrastructure layer share the same end goal but operate on fundamentally different capital schedules and risk profiles.

The capital distribution confirms the asymmetry. Robotic Foundation Models captured 44.9% of disclosed Physical AI funding. General Purpose Robots captured 32.7%. Together, these two categories accounted for over three-quarters of all disclosed capital. Enabling categories such as simulation, developer tools, and fleet management remained marginal by comparison. The structural implication is that Physical AI is not a diversified ecosystem of evenly weighted sub-sectors. It is a landscape driven by concentrated capital in companies whose ambition is platform-level intelligence. Enabling layers, despite being essential to the ecosystem's function, are systematically underfunded. Understanding this asymmetry is the first requirement for mapping the landscape.


Theme I: Foundation Models and World Models

The highest-premium segment of the Physical AI stack is the intelligence layer. Robotic foundation models are large-scale, pre-trained AI systems designed to generalise across robot bodies and tasks, learning from diverse data sources including internet video, simulation, and physical demonstrations. They represent the Physical AI equivalent of what GPT-class models did for language: models trained to control physical bodies instead of generating text. World models are a closely related but architecturally distinct subcategory, neural networks trained on vast quantities of video and sensor data that learn the physics of the world through observation, enabling robots to predict environmental dynamics, plan multi-step actions, and recover from errors before a physical mistake occurs.

The structural logic behind this layer rests on a simple premise. In the same way that LLM foundation models captured value across hundreds of downstream software applications, a cross-embodiment robotic foundation model can capture value across multiple hardware platforms and deployment environments. The companies that build this intelligence layer do not need to manufacture robots. They need to own the data loop that continuously improves model performance with every real-world deployment hour, a compounding mechanism that Bessemer Venture Partners identifies as the structural advantage separating platform-level intelligence companies from model-architecture competitors (Bessemer, 2026).

The technology signal is unambiguous. Scaling laws that reliably improved LLM performance with more data and compute have now been demonstrated in robotics. Foundation models trained on diverse physical interaction data improve predictably as pretraining data size increases. The same architectural principle that drove the LLM era, that a general-purpose model trained broadly outperforms a collection of narrow task-specific systems, is being confirmed in the physical domain. Models trained on internet-scale video and real-world teleoperation data can now handle unfamiliar environments and new tasks without per-environment retraining, a capability that traditional hand-coded robot control systems cannot achieve.

Lens

  • Market Sizing: Grand View Research projects the broader Physical AI market to reach USD 960.38 billion by 2033 at 36.1% CAGR. Investment in world models and VLA models grew by a factor of five from 2024 to 2025, from USD 1.4 billion to USD 6.9 billion (CB Insights).
  • Capital Concentration: Robotic Foundation Models captured 44.9% of all disclosed Physical AI funding with only 28.1% of deal volume. The median Physical AI round across the ecosystem is USD 112.5 million, and the majority of disclosed rounds exceeded USD 50 million.
  • Structural Dynamics: Cross-embodiment intelligence is the principal architectural advantage in this layer. The critical differentiator is a proprietary data flywheel: real-world deployment data that continuously improves performance, creating a compounding mechanism that cannot be replicated through model architecture alone (Bessemer Venture Partners, 2026).

Key Players

  • Skild AI builds a scalable robotics foundation model trained on internet-scale video data, with reported commercial revenue generation within months of product launch.
  • Physical Intelligence develops the π0 series of robot foundation models, demonstrating cross-environment task generalisation without per-environment retraining.
  • FieldAI builds field foundation models for robots operating in complex, unstructured physical environments.
  • Generalist AI produces foundation models for embodied intelligence, with published results confirming that robotics training follows the same scaling law dynamics that defined the LLM era.
  • General Intuition trains AI models on large-scale video game and action data rather than physical robot demonstrations.
  • World Labs, founded by Fei-Fei Li, builds 3D world models that enable robots to understand and navigate physical spaces from video data alone, representing the most capitalised pure-play world model company.

Theme II: General Purpose Robots and Humanoids

The embodiment layer combines hardware, software, and deployment infrastructure into integrated physical platforms. Unlike the intelligence layer, where companies focus on model capability without owning hardware, embodiment companies must solve manufacturing, supply chain, field reliability, and deployment economics simultaneously. This means higher upfront costs, longer development cycles, and a success condition that depends on physical deployment density rather than model accuracy alone.

The market draws a meaningful distinction between general purpose robots and humanoids. General purpose robots are integrated hardware and software platforms deployed in specific operational contexts: warehouses, factories, logistics centres. Humanoid robots optimise for the human form factor, reflecting the view that environments built for humans require human-shaped machines. Capital data shows a structural preference for full-stack platforms with demonstrated deployment capability over form-factor bets that depend on cost breakthroughs that have not yet materialised at scale.

Hardware cost compression is reshaping where value accrues in this layer. The long-run cost trajectory for industrial robotics is sharply downward, driven by commoditised sensors, standardised actuators, and manufacturing scale. This compression favours companies whose competitive advantage resides in proprietary data pipelines and deployment density rather than differentiated hardware. A factory floor generating continuous training data from hundreds of deployed robots creates an operational moat that component cost alone cannot breach. The structural consequence is that value in the embodiment layer migrates over time from the manufacturer to the operator with the largest deployment footprint.

Lens

  • Market Sizing: Goldman Sachs projects the robotics market at USD 38 billion by 2035, a forecast revised upward sixfold in a single year. The sector has attracted fewer scaled companies than software: 745 software companies have raised over USD 30 million in the past five years, compared to 42 robotics companies (Bessemer Venture Partners).
  • Capital Concentration: General Purpose Robots and Humanoid Robots together accounted for 49.8% of all disclosed Physical AI capital. General Purpose Robots commanded a meaningfully higher capital-per-platform metric than Humanoid Robots, reflecting the structural advantage of demonstrated deployment density.
  • Deployment Economics: Hardware cost compression is unlocking deployment scenarios that were economically infeasible five years ago. The mid-market manufacturing layer, comprising over 200,000 facilities in the United States alone, has no systematic access to physical AI capabilities. Bridging frontier model intelligence with mid-market deployment economics represents the largest structural gap in the embodiment landscape.

Key Players

  • Mind Robotics, the Rivian spinout, builds AI-powered industrial robots trained on proprietary manufacturing data, exemplifying the deployment-data-as-moat thesis.
  • NEURA Robotics develops cognitive robots and the Neuraverse physical AI platform, positioning as Europe's integrated full-stack embodiment bet.
  • Figure AI is the most capitalised pure-play humanoid company, designing autonomous general-purpose humanoid robots for real-world work environments.
  • Standard Bots builds AI-native industrial robot arms for autonomous factory automation.
  • Agility Robotics deploys Digit humanoid robots in warehouse environments, representing the closest the market has to production-scale commercial deployment.
  • LimX Dynamics develops legged and humanoid robot platforms, representing the Asia-Pacific hardware pipeline.
  • Rhoda AI bridges foundation model intelligence and real-world industrial production environments, focusing on the deployment layer between model capability and physical operations.
  • Dyna Robotics develops robotic foundation models targeting general-purpose deployment in commercial environments.

Theme III: Simulation, Data and Developer Infrastructure

The Physical AI ecosystem depends on an enabling infrastructure layer that encompasses simulation platforms, synthetic data pipelines, teleoperation infrastructure, developer tooling, and fleet observability. This layer is structurally essential to every other segment: foundation models cannot train without simulation and data, and embodiment platforms cannot deploy without debugging and monitoring tools. Yet it is the most systematically undercapitalised layer in the entire stack, accounting for less than 6% of disclosed Physical AI funding despite being a dependency for every model company and robot manufacturer.

The structural data problem is the defining constraint. Where LLMs bootstrapped on trillions of freely available internet text tokens, Physical AI has no equivalent training corpus. Total global robot manipulation data is estimated at roughly 300,000 hours, compared to approximately 1 billion hours of internet video. Simulation platforms, synthetic data generators, and teleoperation fleets are the primary mechanisms for closing this gap, generating training data at a scale impossible to achieve through physical demonstration alone. Bessemer Venture Partners estimates aggregate robotic data costs across the industry will exceed USD 3 billion, spanning teleoperation, egocentric video, simulation, and physical demonstration collection.

Unlike the LLM market, where developer tooling is mature and well-capitalised, Physical AI developers operate in a tooling environment that is years behind their software counterparts. Data management, debugging, observability, and fleet monitoring platforms for physical AI systems receive a fraction of the capital flowing to model companies and hardware platforms, even though every deployed robot fleet generates a continuous demand for exactly this infrastructure.

The simulation and digital twin segment of this layer was mapped in detail in our previous analysis, Market Deep Dive: Industrial Simulation, which examined the architecture shift from physics-grounded simulation to probabilistic world models and agentic orchestration platforms. That analysis found the digital twin market projected to reach USD 328.51 billion by 2033 at 31.1% CAGR.

Lens

  • Market Sizing: The simulation and digital twin market is projected to reach USD 328.51 billion by 2033 at 31.1% CAGR (Grand View Research), with the AI-powered subset reaching USD 15.24 billion by 2032 at 32.6% CAGR (MarketsandMarkets). Aggregate robotic data costs across the industry are projected to exceed USD 3 billion (Bessemer Venture Partners).
  • Capital Concentration: Enabling layers together accounted for under 6% of disclosed Physical AI capital. Developer tools captured roughly 1.7%, simulation training captured 0.7%, and fleet management captured 0.8%.
  • The Sim-to-Real Challenge: Robots achieving 95% task success in simulation typically achieve only 30% to 50% in real environments. Uncontrolled lighting, surface variation, and unexpected objects create a reliability gap that no simulation platform has closed to date.

Key Players

  • NVIDIA occupies the most structurally defensible infrastructure position. The Cosmos world model platform and the Isaac simulation ecosystem provide foundational infrastructure on which much of the Physical AI ecosystem builds.
  • Scale AI provides synthetic data infrastructure for physical AI model training, with major technology companies making significant direct investments in the data pipeline layer.
  • Foxglove builds a data platform for robotics and autonomous system developers, providing debugging, visualisation, and data management tools.
  • Vention provides a cloud robotics platform enabling manufacturers to design, deploy, and manage robotic workcells.
  • Flexion Robotics develops reinforcement learning and sim-to-real training platforms specifically for humanoid robot intelligence.
  • Zeromatter develops synthetic data generation infrastructure, challenging the industry assumption that teleoperation alone can supply sufficient training diversity.
  • Antioch builds cloud simulation tools for testing and validating robotics autonomy.
  • The broader simulation ecosystem includes Duality AI, Parallel Domain, Voxel51, and Genesis AI, each building physics engines, data labelling, and synthetic data pipelines.

Structural Constraints

The Physical AI ecosystem operates under a set of structural constraints that define its timeline and risk profile. These cut across all themes and determine which architectural approaches are commercially viable at scale.

The sim-to-real gap is the central unsolved engineering problem. Robots achieving 95% task success in simulation typically manage only 30% to 50% in real environments. The physical world is adversarial in ways that simulation cannot fully capture: uncontrolled lighting, surface variation, unexpected objects, and human unpredictability. NVIDIA's Cosmos, Meta's video prediction architectures, and open-source physics engines are attacking this problem, but no standardised bridge has emerged. Closing this gap is the precondition for generalised deployment.

Data scarcity is a structural bottleneck that capital alone cannot immediately solve. The LLM era bootstrapped on freely available internet text. Physical AI has no equivalent corpus. An estimated 300,000 hours of global robot manipulation data exists, compared to roughly 1 billion hours of internet video and 300 trillion tokens of text. Every company in the stack depends on data generation infrastructure that does not yet exist at the required scale.

Unit economics for physical hardware remain challenging. Most humanoid robots today cost hundreds of thousands of dollars while offering battery life in the range of two to five hours. Supply chain concentration is a concurrent risk, with an estimated 90% of humanoid robot components manufactured in China. Manufacturing scale, safety certification timelines, and component costs keep deployment economics uncertain for all but the most capitalised players.

Talent concentration compounds these risks. Among US robotics companies that have achieved growth-stage scale, 43% of founders hold PhDs. Nearly half come from just four institutions: Stanford, MIT, Berkeley, and Carnegie Mellon. Robotics does not have a broad-based talent ecosystem. It is a narrow pipeline producing a small number of exceptionally capable people, and those people are already concentrated in a small number of companies. This limits the rate at which new fundable teams can form and scale.


Takeaways

  • Physical AI investment is not distributed evenly across the stack. It is concentrated in the intelligence layer, where Robotic Foundation Models and General Purpose Robots captured 77.6% of disclosed funding. Enabling layers remain systematically undercapitalised, creating an asymmetry that represents both a risk and an opportunity for structured ecosystem mapping.

  • Cross-embodiment intelligence is the central structural dynamic of the current Physical AI cycle. Companies that own the most effective data flywheels strengthen their position as deployment scales. Hardware, by contrast, is subject to commoditisation pressure that shifts value over time from differentiated manufacturing toward deployment density and proprietary operational data.

  • The Physical AI stack separates into three distinct layers, intelligence, embodiment, and enabling infrastructure, each with fundamentally different capital requirements, technology maturity, and competitive timelines. The enabling infrastructure layer, despite being a structural dependency for every other segment, remains the most systematically undercapitalised portion of the landscape.


Sources & Citations

Nami Venture Partners