AWS published a technical guide on September 4 showing developers how to build a Physical AI model factory: the continuous pipeline of synthetic data generation, model post-training, and closed-loop evaluation that robotics and autonomous vehicle teams need, running on Nvidia's Cosmos 3 world foundation model. The guide, written by AWS solutions architects Nathan Arnold and Eric Saleh, runs that pipeline on a persistent Amazon SageMaker HyperPod cluster orchestrated with Amazon Elastic Kubernetes Service (EKS), and argues that GPU goodput (useful pipeline progress per reserved GPU hour) is the metric that actually determines whether the setup pays for itself, not the peak throughput of any single job.

Cosmos 3 itself is not new. Nvidia launched it on June 1, 2026 at GTC Taipei, positioning it as the first fully open omnimodel that treats video, image, sound, and robot action as a single token stream. Nvidia CEO Jensen Huang framed the underlying advances at the time as heralding "the big bang of physical AI." AWS's new post is effectively the infrastructure playbook for running that model at production scale on its own cloud.

Why One Training Job Isn't Enough

A robot or self-driving system cannot be built with a single fine-tuning run, according to AWS's guide. Instead, teams need a loop: ingest and curate real-world sensor data, generate synthetic data to cover scenarios too rare or dangerous to collect safely, post-train perception and policy models on the combined corpus, evaluate the result in closed-loop simulation, and feed whatever fails back into the next round of data generation. AWS calls this cycle the Physical AI "flywheel." Its core argument is that provisioning a separate GPU pool for each stage of that loop wastes capacity, since teams end up standing infrastructure up and tearing it down between stages instead of keeping it busy.

What Makes Cosmos 3 Different

Cosmos 3 uses a Mixture-of-Transformers design that pairs a reasoning "expert" with a generation "expert" joined at every layer, rather than bolting a video generator onto a separate language model. Because one architecture can run in three modes (a world model that generates synthetic video, an action labeler that converts raw video into labeled training data, and a deployable robot policy), the same model family covers every stage of the flywheel. That is what lets AWS collapse three separate GPU pools into one shared, time-shared cluster.

Nvidia ships Cosmos 3 in three size tiers, each aimed at a different job in the pipeline:

Tier Parameters Backbone Role in the factory
Cosmos3-Nano ~16B Dense 8B Qwen3-VL Fully fine-tuned into the deployable robot policy
Cosmos3-Super ~64B Dense 32B Qwen3-VL LoRA-adapted synthetic-data teacher
Cosmos3-Edge ~4B ~2B, trained from scratch On-device deployment, benchmarked on Jetson Thor and Orin

Nvidia released Cosmos 3 under the Linux Foundation's OpenMDW-1.1 license, and AWS's guide, along with the accompanying awsome-distributed-ai GitHub repository, builds directly on Nvidia's own cosmos-framework training stack without modifying it.

Why SageMaker HyperPod

AWS's pitch is that the demands of running Cosmos 3 continuously map onto specific properties of Amazon SageMaker HyperPod on EKS: one persistent GPU pool that all three pipeline stages share, a single Amazon FSx for Lustre storage layer reachable over Elastic Fabric Adapter (EFA) so generated data, training checkpoints, and evaluation traffic never have to move between clusters, automatic detection and replacement of failed nodes with job auto-resume from the last checkpoint, and optional task governance (built on the open source Kueue scheduler) for teams running several robot or vehicle programs against the same reserved capacity at once.

AWS ran its own tests on p5en.48xlarge nodes, each with eight Nvidia H200 GPUs. Scaling the 64B Cosmos3-Super LoRA workload from one to four nodes (8 to 32 GPUs), AWS reported the run held between roughly 0.97 and 0.99 of linear scaling efficiency, with per-step time varying by about 3% across that range. Measured Model FLOPs Utilization landed near 0.50 for the compute-heavy Super workload and near 0.24 for the lighter Cosmos3-Nano vision fine-tuning workload. AWS is explicit that these are relative, single-run figures meant to illustrate a methodology, not a benchmark leaderboard.

Why It Matters

This is not a new model launch or a splashy product reveal. It is a reference architecture, and reference architectures like this one have become a recurring genre: AWS previously published a similar developer playbook for building multi-tenant agentic chat apps on Amazon Bedrock, and this Cosmos 3 guide follows the same format applied to a very different workload. Every time a major foundation model ships, the cloud providers that host its training racks put out a guide showing how to run it well on their own infrastructure. That is genuinely useful for the developers who need it, but it is also a proxy fight for GPU market share between the hyperscalers, playing out one blog post at a time.

What is more notable is the framing choice: AWS centers the entire piece on GPU goodput rather than headline throughput numbers, which reflects a real problem physical AI teams face. GPU capacity at this scale, especially the p5en.48xlarge instances this guide is built around, is expensive and frequently supply-constrained, so a team pays for reserved capacity whether or not its pipeline is making progress on it. Treating idle or duplicated capacity as the thing to eliminate, rather than treating peak speed as the goal, is a more honest way to talk about training economics than most vendor content manages. It also fits a broader pattern this year of physical and embodied AI graduating from research demos into production tooling, the same shift reflected in TechCrunch Disrupt 2026 adding a dedicated stage for real-world AI this year.

What to Watch

Watch whether Microsoft Azure or Google Cloud publish comparable Cosmos 3 reference architectures of their own, since Nvidia's open license permits it and the competitive incentive clearly exists. Watch also for autonomous-vehicle-specific results: this guide sets up the pipeline to extend to AV post-training but only demonstrates the robot-policy and vision fine-tuning workloads end to end. Finally, watch capacity availability for p5en.48xlarge instances specifically, since AWS's own guide flags that GPU supply, not raw compute pricing, is the binding constraint for teams trying to actually run this setup.

Key Takeaways

  • AWS published a guide on September 4 for running a continuous Physical AI training pipeline (synthetic data generation, post-training, and closed-loop evaluation) using Nvidia's Cosmos 3 on Amazon SageMaker HyperPod.
  • Cosmos 3, which Nvidia launched in June 2026, ships in three tiers: a 16B Nano policy model, a 64B Super model used as a synthetic-data teacher, and a 4B Edge model for on-device deployment.
  • AWS's own tests on p5en.48xlarge (8x Nvidia H200) nodes showed the 64B workload holding roughly 0.97 to 0.99 of linear scaling efficiency from 1 to 4 nodes.
  • The guide frames GPU goodput, not peak job throughput, as the metric that determines whether a physical AI training pipeline is cost-effective.

FAQ

What is a Physical AI model factory?

It is AWS's term for the continuous pipeline that robotics and autonomous vehicle teams need to run rather than a single training job: ingesting real-world data, generating synthetic data for rare scenarios, post-training perception and policy models, and evaluating the results in closed-loop simulation before feeding failures back into the next round.

What is Nvidia Cosmos 3?

Cosmos 3 is Nvidia's open world foundation model for physical AI, launched June 1, 2026 at GTC Taipei. It uses a Mixture-of-Transformers design that treats video, image, sound, and robot action as one token stream, letting a single model function as a synthetic-data generator, an action labeler, and a deployable robot policy.

Why use Amazon SageMaker HyperPod instead of a standard training job?

AWS's guide argues HyperPod fits Physical AI workloads specifically because they run continuously rather than as a one-off job. HyperPod provides a persistent, health-monitored GPU pool with automatic node replacement and job auto-resume, which AWS says is not necessary for a short, one-shot fine-tuning run but becomes important once a pipeline runs for days across many nodes with statistically frequent hardware failures.

Is Nvidia Cosmos 3 open source?

Nvidia released Cosmos 3 under the Linux Foundation's OpenMDW-1.1 license. AWS's guide and its accompanying GitHub repository build on Nvidia's own cosmos-framework training stack without modifying it.