← Projects

When a Simulation Result Can Count as Evidence

Why a 664 mm slide was an acceptable delivery, and how the worker tied that result to the scene

In this article 5 sections

The block moved about 664 mm down the incline, well beyond the 10 mm comparison limit. The harness still accepted that scene under the production brief. The assignment was to preserve a supplied model, set its mass to 1.25 kg and deliver three friction variants. It did not ask every variant to stay put.

Changing the friction until the block passed would have undone the requested comparison. The displacement was useful evidence about the low-friction variant, not a reason to replace it with a different one.

This follow-up puts the earlier incline experiment behind a bounded harness worker. I wanted the application to receive enough evidence to decide what to do with the result: which scene ran, whether the run completed and its trace supports the reported displacement, and whether the brief makes that displacement a reason to reject.

The same comparison can serve different tasks

The September 24 brief asked for friction values of 0.22, 0.42 and 0.68 on the supplied two-box model. Three attempts delivered nine scenes. The selected checks confirmed the requested parameters and preservation of the base. Displacement was measured as an advisory comparison.

The evaluating assistant ran all nine scenes at 1 ms and 0.5 ms timesteps. Their 18 traces matched direct execution through the same adapter. At friction 0.22, the maximum displacement was 664.272 mm at 1 ms and 663.941 mm at 0.5 ms. All three attempts gave the same respective values; repeated agreement does not make them independent physical measurements.

Displacements in the supplied friction variants. At coefficient 0.22, the block moves about 664 millimetres; at 0.42 and 0.68 it moves about 1 millimetre. Separate scales retain the small movements. All nine deliveries satisfy the requested variant-authoring task; the 10 millimetre behavior comparison is advisory.
Maxima from the retained worker traces, grouped by friction and timestep. Each value was the same across three attempts. The low-friction result exceeds the advisory 10 mm comparison. The brief required those variants, so that observation did not reject them. Friction remains assumed.

The original incline experiment had a different contract: staying within 10 mm was required. There, the complete low-friction trace correctly produced a rejection. The two contracts give the same kind of measurement different consequences because they ask for different work.

An extra-object control tested another boundary. Selected physics-configuration rules passed, but the restricted worker returned unknown because it only covers the two-box structure. A separate base-preservation check rejected the copy under this brief. Without that requirement, an extra object would call for a more capable evaluator; it would not be inherently wrong.

Tie the worker to the scene being accepted

The harness admits the saved USD, copies it into a temporary job directory and launches fixed installed worker code in a separate Python process. The job records the input hash, timestep, duration, and worker and adapter source hashes. The candidate cannot choose an executable or submit Python to run.

Review found that the worker and adapter hashes were recorded without being verified. The fix included in release 0.5 checks the installed worker, adapter and fixed profile before and after execution, then checks their returned identities in the parent process. A missing or changed identity cannot supply accepted evidence. This is a new regression-tested safeguard; the older receipts below do not establish it.

The adapter accepts only the fixed ramp and one rigid block, with limited mass and friction variations. It reads that supported USD structure and writes an explicit MuJoCo model. This is a small CPU experiment using MuJoCo’s Python bindings, not a general USD-to-simulator conversion. Keeping the model unchanged from the earlier experiment made it possible to check whether the worker handoff altered the measurement.

The worker returns its identities and runtime versions, simulated times, block positions, displacement samples and the maximum displacement. The parent first checks that the record belongs to the current job. A low value from yesterday’s scene cannot pass today’s requirement simply because it is below the limit.

It then checks trace length, time spacing and finite values. From the saved positions and ramp angle it recomputes the displacement along the slope and its maximum absolute value. A reported scalar that disagrees with the trajectory is rejected as unusable evidence.

These checks catch stale or inconsistent output from the trusted worker. They cannot establish that an arbitrary remote worker honestly ran a simulator: a fabricated but internally consistent trajectory could pass. The model and CSV trace are retained so a reviewer can inspect what ran and recompute the metric.

Check the handoff against the earlier experiment

Before using the new producer variants, I kept the original 1 kg block and friction values of 0.15 and 0.65 as regression controls. The standalone runner, direct worker execution and integrated harness reproduced these maxima in the recorded environment:

Scroll sideways for more columns.

Assumed friction1 ms step0.5 ms stepOriginal required limit: 10 mm
0.651.064 mm1.045 mmPass.
0.15987.227 mm986.730 mmFail.

The comparison checks the integration. The paths share the adapter and simulator, so agreement does not independently validate the contact model. The earlier article’s trace analysis discusses the ideal calculation and the small high-friction movement that it does not predict.

The new pack also keeps assumed parameters visible. When a control requires measured friction, the passing high-friction trajectory returns insufficient evidence. This prototype has no workflow for admitting and validating a measured friction record. Successful execution cannot supply that missing basis.

A worker failure needs a different response

The caller sets a wall-clock deadline. This small job is also limited to 10,000 simulation steps, with output monitored by the parent and evidence reading capped at 4 MiB. Those limits let the caller stop waiting for an unusable job. They are not isolation from hostile code: the fixed worker runs under the same operating-system user, without a hard memory limit or network sandbox.

I tested deliberately interrupted and malformed runs, along with wrong identities and inconsistent trajectories. These are development faults introduced to test the handoff, not a measured rate of production failures.

Scroll sideways for more columns.

ObservationResult under the original stay-within-10-mm contractCaller action
Complete, matching trajectory exceeds 10 mmREJECTInvestigate the model or requirement; do not silently change contact assumptions.
Complete trajectory stays within 10 mmACCEPT_FOR_USERetain the model assumptions with the simulated result.
Timeout, crash, malformed or oversized outputEVALUATION_ERRORFix or retry the execution. There is no usable behavior measurement.
Wrong input/job identity or inconsistent scalarEVALUATION_ERRORInvestigate the evidence handoff.
Scene outside the two-box adapterINSUFFICIENT_EVIDENCEUse a suitable evaluator.
Measured friction required but absentINSUFFICIENT_EVIDENCESupply the missing parameter basis through a supported workflow.

The timeout control deliberately sets an inadequate deadline. Crash and malformed-output controls run substitute subprocesses through the same parent function; further injected records exercise identity and trajectory checks. The test source keeps those mechanisms separate from the valid simulations.

Reproduce the worker and its controls

Scene Acceptance 0.5 on GitHub includes this restricted block-and-ramp worker and the identity-verification fix described above. The companion preserves the earlier fixed-model experiment and its fault controls. It does not include that fix or the later nine-scene variant study.

The v0.1 companion includes the fixed adapter, installed pack, original frozen inputs, traces, reports and fault tests, with private environment paths removed from the public copy. This snapshot is based on harness commit 33e5721; it is separate from both v0.3.0 and the current 0.5 release. The later nine-scene evaluation and supplemental preservation checks remain in the private evaluation/fresh-producer-v1/ workspace, outside this companion.

After the companion installation, the matrix runs the worker cases and compares them with direct execution:

python evaluation/followups-v1/run_matrix.py /tmp/worker-replay-01
python -m pytest tests/test_followups.py -q

The case folders retain report.json, and completed simulations also have trace.csv, model.xml and a direct-run result where applicable. The test source shows each injected fault. The original experiment remains separately replayable with its pinned article dependencies.

The recorded runs use Python 3.12.14, OpenUSD 25.11, MuJoCo 3.6.0 and NumPy 2.5.3 on macOS arm64. Code and constructed controls were produced with coding-assistant help under my direction. The original configuration comparison uses NVIDIA USD Validation 1.20.0; the new simulation result comes from MuJoCo and the bounded adapter. Neither supplies measured material properties.

The 664 mm result is the case I would keep visible when adding a repair loop. The producer delivered the requested low-friction variant. A loop that kept changing it until the advisory displacement turned green would have made the submission worse for the task we actually gave it.

Disclaimer: The views and opinions expressed in this account are those of my own and do not represent those of my employer, NVIDIA.

← All projects