← Writing

What a Robot Model Still Leaves to the Plant

I originally framed this as “the bottleneck isn’t the model anymore.” That was too broad. Model capability still limits what a robot can do, and a vision-language-action model can itself be the policy that produces its actions. My argument is about the work a plant still owns when it starts from that policy. I would build the data, validation and recovery pipeline now, because a better checkpoint does not establish whether a particular cell can keep meeting its production requirements.

If you know me, you already know the use-case I keep coming back to: automotive parts kitting. An order flashes on a screen; a person reads it, picks the right part from one of a wall of racks, drops it in the kit tote, then the next part, and the next (five or six picks) until the full bill of materials for that kit is in the tray. The tray rides a conveyor to the assembly line, just-in-sequence. It’s among the most common material-flow patterns in discrete manufacturing. Repetitive, and (because the cell’s already monitored) on camera all day. The tempting question: can a robot learn it just by watching the people who already do it?

A parts-kitting cell: an order screen feeds a wide pick face of many storage racks with variable contents. The task transitions from a human picker to a robot that verifies each part before making five to six picks into a kit tote, which a conveyor carries as a line of filled totes to an assembly station just-in-sequence.
Automotive parts kitting: verified picks from a wall of racks build a kit tote, which rides the conveyor to assembly just-in-sequence, with the task shifting from human to robot.

In this cell, recognizing the parts, interpreting the order and moving the gripper are related jobs. A vision-language-action policy takes observations and an instruction and produces robot actions. NVIDIA’s GR00T architecture, for example, combines image, language and robot-state inputs with an action-producing network. Developers can adapt a pretrained policy to a robot and task instead of starting every skill from scratch.

That policy still sits inside a larger system. The order service has to supply the right bill of materials. The robot’s controller has to execute the requested motion within the cell’s limits. Verification has to establish that the part in the tote matches the order. A model that describes a grasp or predicts a plausible next view does not, by that result alone, establish that the kit passed QA. I would test the policy and those surrounding checks together.

The first job is establishing what the pretrained policy can already do on this machine and where it needs adaptation. Human video can contribute task knowledge, but it does not contain this gripper’s joint commands. A failed grasp may call for robot demonstrations, a changed fixture or more representative training conditions; I would use the failure to decide. Data collected on the hardware and data generated in simulation both need validation against the real task, especially when contact decides whether a part seats or jams. The training pipeline is how the team keeps that work repeatable as the policy changes.

The second job is responding when the work changes, and drift here isn’t exotic. A design change goes into pilot production, and for a few weeks a slightly different part runs down the line. Or a component is short, so an approved alternate is substituted for a while: same function, a different part number, close to the original but not the same. A person may absorb that change with little difficulty. The team still has to establish whether the robot can do the same. And it stacks on the mundane stuff: bins get restocked in the wrong slot, SKUs get added, orders change this morning.

The robot has to verify what it picks and check the finished kit against the order. A failed check needs a defined response, and the recorded failure should inform the next change to the cell or policy. Retraining is one possible response; some changes only need revalidation, a configuration update or a correction to the source data. I would make that decision from the failure evidence rather than automatically retrain after every mismatch.

You can’t let a robot freely explore on a live just-in-sequence line to learn. A mis-pick can stop the line or send the wrong kit to the assembly station. Changes need a place to be tested before they reach production, with an operating boundary around the policy that defines what happens when a check fails.

The strongest argument for waiting is that a more capable generalist policy could reduce how much task-specific data and tuning this cell needs. I expect that improvement to matter. It could change the economics enough to make an impractical task worth attempting. It still leaves the plant responsible for the order integration, acceptance criteria and response to a wrong pick. Building those parts now also gives the team a way to evaluate the next model against the same production requirements.

What would change my view of the adaptation burden is a kitting cell that absorbs a new SKU from a handful of human runs and still holds up a week later, with no manual data repair between them. I would want to see the mis-picks, interventions and recovery time as well as the successful kits. Those are the results I would use to decide how much of the pipeline the next cell needs.

Disclaimer: The views and opinions expressed in this account are those of my own and do not represent those of my employer, NVIDIA.

← All writing