What it takes to make AI hold up in the physical world.
Understand what to test before relying on a robot policy, digital twin or agent workflow. Find practical explanations, measured comparisons and deployment checks for the decisions between a working demo and a dependable system.
First reads
Start with a decision you need to make.
A VLA Benchmark Is Not a Commissioning Package
What to validate, monitor and plan for before deploying a learned robot policy.
Reconstruction qualityWhat I Learned from 24 Gaussian Splat Jobs
How to compare training time, Gaussian budgets and image quality before spending more GPU time.
Reliable agent workflowsI Took the Order Away From the Agent. Then I Learned Order Wasn't Enough.
What to check before retrying an interrupted agent run.
Browse by topic
Choose an interest area, then follow the reading path that fits what you came for.
Learning, deployment and commissioning
Understand how capture, training and validation fit together, and what a learned robot policy still needs before deployment.
Simulation & Digital Twins · 4 reading paths · 28 notesFrom evidence to a usable twin
Choose the representation, simulation model and application architecture that fit the decision your digital twin needs to support.
Agentic Systems · 3 reading paths · 16 notesGrounding, harnesses and reliable execution
Decide when better context, a structured workflow or stronger runtime checks will make an agent more reliable.
Platform & Industrial Systems · 1 reading path · 13 notesFrom compute to operation
Understand where hardware, Omniverse and shared models fit in an industrial system, and which layer owns each capability.
Recent
Six readable labels exposed why a pixel comparison must retain its regions and tolerances. A later detailed-brief study found wrong source pixels despite passing geometry and motion checks.
A connection check caught deliberately broken motion, while three supplied mechanisms passed its samples. A valid speed change exposed an extra requirement in the evaluator, and continuous attachment remained unresolved.
A low-friction variant exceeded a behavior threshold and still satisfied its production brief. The worker validates the input identity and trajectory before the harness decides whether that measurement calls for repair.
I built a common acceptance workflow for my own 3D jobs, with the brief defining the required outcome and packs supplying selected checks. A panel test and a 24-scene study show what that adds, where deliveries still fail general checks, and which questions remain open.
Astra built an animated crank-slider with materials, editable USD and an interactive browser delivery. This recorded example shows what the agent handled, what it corrected and what I would still check before using the result.
A dependency reader missed a shader-authored texture path and falsely accepted three cases. The corrected material checks resolve bindings and missing files; a separate decode probe shows what file presence still leaves open.
Controlled tests expose a wrong midpoint, an excursion between samples and a playback duration missing from the contract. A new timing pack rejects the clock-only change while accepting a correctly rescaled animation.
A runnable project combines selected NVIDIA USD configuration checks with a separate MuJoCo behavior experiment. Two configurations pass, but only one meets a 10 mm displacement limit. The walkthrough shows the inputs, evaluation settings, traces and unresolved physical evidence.
I built SplatStage around editing, export and composition workflows. Starting again, I would test delegating that production job to a runtime agent. The question is which development work that avoids, and which tools, checks and user-facing features I would still maintain.
Four controlled cases expose an incompatible pallet-move request and a checker that never tested the destination. A feasible layout can pass while leaving the original request unresolved; the code and saved scenes are available to reproduce.
Five Astra probes produced no hallucination example. One correctly left contact, mass and friction unresolved. Together with a published object-grounding failure, that result shows why spatial evaluation needs supported, contradicted and unknown outcomes.
A digital twin needs a surrogate only when repeated prediction creates value and the solver misses that decision's clock. Ventilation makes the case; fast energy models show why the rule is not universal. Part 1 of a five-part series.
SurrogateLab is the executable workbench behind the SurrogateGate decision framework. It compares classical reduced models, neural methods and PhysicsNeMo implementations across contrasting 2D and 3D problems. Part 2 of a five-part series.
FNO led three 2D tests; POD-NN led two and shared one top score. The useful result was what projection error, regression capacity and data support changed in the next experiment. Part 3 of a five-part series.
A 16-case test made POD+GP look better than a PhysicsNeMo FNO. A common 200-case test put POD-NN first and left FNO and GP 0.03 points apart. The evaluation design changed the conclusion. Part 4 of a five-part series.
A 2D room surrogate answers repeated field queries quickly, but the whole-field winner is not the winner on every comfort quantity. This is the gap between a fast model and a decision-ready building loop. Part 5 of a five-part series.
Years of MES work taught me where automation stops: at tasks whose decisive state cannot be written down or seen. Tactile input earns its place when it reveals that missing state; contact alone is not enough.
Anchor stopped my agent from skipping declared steps. Ratchet came from the next question: after an interrupted or overlapping run, could I prove what happened and restart without making it worse?
A technical look at the small execution runtime I built for overlapping and interrupted agent workflows: ownership, effect recovery, verified completion and the limits of a local alpha.
A Gaussian splat can reproduce a place convincingly while knowing almost nothing about its identity, scale or behavior. This article separates reconstruction, visual readiness, operational readiness and simulation readiness—and defines the boundary SplatStage is designed to cross.
The editor showed a cleaned scene while the exporter still read the original checkpoint. SplatStage makes the selected edit version the export input and checks it again through OpenUSD readback. The build choices behind that handoff.
One Garden scene made the promises and gaps in SplatStage measurable: 2.07 million Gaussians removed, an edited OpenUSD particle field, three composed assets, a dependency-complete stage—and no measured scale, colliders or external runtime proof.
Every digital twin needs a usable digital starting point. CAD, LiDAR, photogrammetry and Gaussian splatting preserve different kinds of truth, so the representation should follow the decision. Part 1 of a three-part series.
A finished reconstruction can still be hard to diagnose. ReconStudio keeps camera-solve evidence, run settings and exported artifacts in one job record, so a user can inspect what happened before trusting the result. Part 2 of a three-part series.
I built ReconStudio, then used 24 job records across six scenes to test its evidence. Metric bugs, repeated runs and paired-frame comparisons changed what I could claim about training, capture and GPU capacity. Part 3 of a three-part series.
A refinery valve operation shows why robot capture should begin with the task: preserve the motion, contact, timing and measurements the policy will need.
Robot foundation models are getting better at handling situations they were not shown. That makes the old commissioning questions more important: where is the model valid, how do I know it has left that range, and what happens next?
A robot can finish the task and still behave unsafely along the way. That distinction makes this benchmark worth reading.
A repeatable agent workflow kept skipping different steps on different runs. The fix wasn't a stronger prompt; it was moving the plan into a DAG and letting deterministic code control sequence and verification.
I choose hardware by the work it runs, the memory it needs and what has to move between devices. A SplatStage run shows why the GPU model alone was a poor guide to sizing the machine.
A month ago I drew a five-stage loop for a robot training center and admitted most of it was an educated guess I hadn't tested. So I built the loop as a real, orchestrated pipeline on one GPU and carried two tasks around it — a pole that balances by trial and error, and a Franka arm that learns to stack cubes by copying demonstrations. This is the environment, the stack, the architecture, and what actually ran. Part 1 of two.
Running the training loop for real surfaced one concept the diagram doesn't capture: the same five stages take two different shapes depending on how the robot learns — by trial and error (reinforcement learning) or by copying demonstrations (imitation). This part is the mechanism behind that fork, grounded in the two tasks I ran, plus the open questions the build left unproven. Part 2 of two.
Before changing model weights, I would identify the gap: missing domain knowledge, unreliable workflow behavior or a deployment constraint. A pharma recipe project, an agent workflow and an edge counterexample lead to different first interventions.
MES was never really about running the machines. It was about improving the operation, and it was limited by how much had to be modeled by hand and decided by people. Physical and agentic AI can change that, but only if the plant's systems are joined through shared operational context.
The useful signal is not that world models are solved. It is that investors and research leaders increasingly see predictive models of physical dynamics as a distinct bottleneck worth funding.
Rendering a USD scene in the browser worked. Keeping edits and rendered state together was harder: separate stages forced reloads and left some session changes invisible to inspection. What I would check before choosing the same architecture.
A pretrained robot policy can supply useful action skills. An automotive parts-kitting cell still needs a way to adapt those skills, verify each kit and respond when parts or conditions change. That is the pipeline I would build now.
A gas-leak prototype showed me the cost of building inside Kit. Standalone Omniverse libraries make the application framework a choice, including for some interactive tools.
Production use no longer requires an NVIDIA AI Enterprise subscription. The useful distinction is between permission to deploy, the license terms that still apply and the support your team needs.
A request to scope a robot training center became a five-stage working model. I use one task to explain the handoffs, the tool choices and what still needs testing before pipeline reuse becomes a business or deployment claim.
From the industrial systems I know, rendering is only one part of Omniverse. The harder job is connecting source data, giving teams a shared context, and reusing it for visibility, simulation, optimization and Physical AI.