← Writing

A Robot Training Center Looks Like Ten Projects. It's One Loop.

In this article 7 sections

A few weeks ago someone asked me a simple-sounding question: if you wanted to stand up a training center for industrial robots from scratch, how would you do it? Not buy a robot for one job, but build the place where a capable, general-purpose robot gets trained into the specific thing you need, again and again. My first reaction wasn’t a plan. It was a small wave of panic. Start listing what such a place touches, capturing how people do the work, building simulations, generating data, adapting large models, proving the results are safe, running them on real robots, and it stops looking like a project. It looks like ten projects, each with its own stack, its own experts, and its own way of failing. I couldn’t fit them into a sentence that held together, let alone a sequence.

I read and talked to people who had built parts of this. Organizing the work around one task helped: capture a demonstration, build and vary the simulation, train a policy, validate it, then deploy it and bring the failures back into the next round. That gave me five stages where I could make a tool choice instead of trying to choose the entire stack at once.

Consider a robot learning to pick a part and place it in a tote. A demonstration alone does not give me a deployable skill. I need a way to vary the situation, train against those examples, check what the policy actually does and bring failed attempts back into the next run. The loop is a working model for organizing that effort. It does not establish that every task needs demonstrations or that a training center will pay for itself.

The examples below lean toward NVIDIA because that is the stack I have used most. I have not run a comparison against the alternatives listed alongside them. My first question is what crosses each stage boundary; the tool choice comes after that.

A cycle of five stages drawn clockwise: Capture at the top, Simulate and Synthesize upper right, Train lower right, Validate lower left, and Deploy and Improve upper left, with a return arrow from Deploy back to Capture closing the loop. Stages 2 to 4 are shaded as the reusable core; Capture and Deploy have dashed borders as the per-customer edges.
The proposed five-stage loop. The return arrow is work to implement and test. Shading the middle as a core worth owning records my initial ownership hypothesis, not a demonstrated business advantage; the data-ownership counterargument remains below.

Capture: record what the robot needs to learn

For the part-and-tote example, a useful demonstration needs observations, actions and the resulting response: where the part was, what the robot was told to do, and whether it actually placed the part. Contact, timing and failed attempts can matter as much as the successful motion. The capture article works through those requirements on an industrial maintenance task.

A teleoperated demonstration can include robot commands. A video of a person doing the job usually does not. Human motion still has to be mapped into a representation the robot and training method can use. I would not treat those two sources as interchangeable just because both show the task.

Projects such as Mobile ALOHA, LeRobot and the Universal Manipulation Interface provide different starting points. Before choosing a capture device, I would specify the signals and variation that the task needs.

Simulate and synthesize: vary the task without losing it

The next stage needs a model in which the robot can attempt the task. For picking a part, that includes the robot, gripper, object and contact surfaces. A convincing background is not enough: the representation must carry the truth the use case needs.

I would start with Omniverse, Isaac Sim and Isaac Lab because that is the route I know. Alternatives such as MuJoCo and ManiSkill offer other simulation paths. The question is whether the chosen environment can represent the behavior being learned and checked.

Synthesis adds variations that would be expensive to capture separately: different part positions, views or other task conditions. Those variations need a reason. More examples from an inaccurate simulator do not by themselves establish better real-world behavior. My later loop build used Isaac Lab Mimic to expand a small demonstration set; Cosmos was a possible additional source, not part of that measured run.

Train: choose how this task gets its learning signal

For a task with usable demonstrations, behavior cloning is one starting point. For a task with a useful reward, reinforcement learning may be another. Adapting a pretrained policy such as GR00T, OpenVLA or π0 is a further option, with its own data and interface requirements.

I originally leaned toward adapting a foundation model. The simpler tests I ran later used reinforcement learning for Cartpole and behavior cloning for cube stacking. They were enough to test the loop mechanics; they did not compare foundation-model adaptation. That is the same discipline I argue for in the fine-tuning article: choose the intervention for the gap in the task, rather than assume the largest training exercise is the right starting point.

The method also changes the earlier stages. My Cartpole run needed no demonstrations, so Capture was explicitly skipped. The imitation path depended on them. The five stages give me common handoffs without requiring identical work inside each one.

Validate: check behavior and keep its limits visible

The part reaching the tote once is not enough to tell me where the policy works. I would test the variations the task requires, retain failures and record which conditions were actually evaluated. Simulation gives a repeatable place to begin, but a good score there does not establish safety or reliable behavior on a real line.

SIMPLER examines the relationship between simulated and real evaluation, while LIBERO and CALVIN provide benchmark tasks. They help answer specific evaluation questions. The integrator still has to establish the operating conditions, limits and response for the actual robot and cell.

This remains the least settled part of my training-center model. I can describe the checks I want; I cannot use a simulation score as a production-safety certificate.

Deploy and improve: build the return path

Deployment asks whether the policy can run at the task’s required rate on the target hardware, and what the system does when it fails or hesitates. I use ROS 2 and the Isaac ROS ecosystem as starting points, but middleware does not settle those operating decisions.

For the tote task, the return path would keep the failed placement, relevant observations and any recovery so the next training run can use them. That is an intended feedback loop, not evidence that repeated training will automatically improve the policy. In the later build, I implemented a file-based hint to the next run but did not demonstrate compounding improvement or deployment on physical hardware.

Compute placement runs underneath every stage: collection, simulation, training and on-robot inference need not run on the same machine. I would make those constraints explicit when scoping the first task.

Choose tools at the stage boundary

These are starting options, not a tested ranking or a requirement to adopt one complete stack.

Scroll sideways for more columns.

StageQuestion before selecting a toolStarting options
CaptureWhich observations, commands and responses must the episode preserve?Isaac teleoperation, ALOHA, LeRobot, UMI
Simulate and synthesizeWhich task conditions need a faithful model and useful variation?Isaac Sim and Lab, MuJoCo, ManiSkill; a separate generation path where needed
TrainIs the task easier to demonstrate or score, and is adaptation justified?Behavior cloning, reinforcement learning, or pretrained-policy adaptation
ValidateWhat does success mean under the declared conditions, and what remains untested?Task rollouts, suitable benchmark cases and cell-specific evaluation
Deploy and improveCan the target meet the response requirement and preserve failures for review?ROS 2, Isaac ROS and a defined feedback path

Separating these choices can make components replaceable. It does not mean that mixing more simulators or tools directly improves generalization. That requires evidence about the data, variation and learned behavior, not just a more varied architecture.

What I would own, and what I still need to test

My initial bet was to invest in the reusable middle: simulation, training and evaluation infrastructure that could carry more than one task. The later build showed that one pipeline and report contract could carry two different learning methods. It did not establish how much reuse would remain across customers or whether that infrastructure would be a defensible business.

The strongest counterargument is that task data matters more than the reusable plumbing. If models and simulators are widely available, the scarce asset may be the demonstrations and field failures that a particular site can collect. That is why I separate two decisions: the capture rig I would rent, the captured data I would guard. I have not run the comparison that would settle where the investment pays off.

I would begin with one real task and specify what each stage must produce and what the next stage must check. The build note records my first test with Cartpole and Franka, and the follow-up explains how their learning methods changed the work. Before expanding that into a training center, I would want the return path to improve a subsequent run and the resulting skill evaluated on the intended hardware.

Disclaimer: The views and opinions expressed in this account are those of my own and do not represent those of my employer, NVIDIA.

← All writing