← Projects
Accepting Agent-Generated 3D Part 1 of 5

Building an Acceptance Harness for Agent-Generated 3D

Working backward from a brief to the checks and evidence a 3D delivery needs

In this article 7 sections

When a 3D job is delivered, I want to know: is it ready for its intended use, and what evidence supports that decision? “The scene passed” was hiding too much. A file could open, satisfy the general delivery rules and still miss a dimension in the brief. A label could have the right source pixels while nobody had checked whether it was readable on the object.

I built this for my own 3D work, starting with the outcomes I needed and working backward to the checks that could support them. That brought geometry, materials, motion, rendered views and some simulation into one acceptance workflow. I spent time on the parts that mattered to those jobs, so the depth varies across them. A different use may need much more from one area and nothing from another.

OpenUSD validators and NVIDIA’s USD validation library already provide extensible ways to check assets. This harness uses selected rules from usd-validation-nvidia==1.20.0 and reports their findings alongside a job’s requirements. Where a task needs a specialist validator or simulator, I would use that through a suitable pack. Bringing the results together does not make each check as deep as a dedicated tool.

Version scope: Scene Acceptance 0.5 includes the interpreter, optional judge and application CLI described here. The studies used local revision 52e412f, tested October 3, 2026; their raw packets remain private and retain their original policy. Later 0.5 fixes have separate regression tests. The public v0.3 release supports the tray replay below; the motion article links the public timing fix.

I asked my coordinating assistant to implement the harness and compare no brief, increasingly detailed briefs and optional model review. The assistant authored the synthetic briefs and evaluated the results. These are exploratory studies, not an independent reliability benchmark.

One panel, a width failure and two missing views

The interpreter test used an existing panel: seven prims, including one mesh and one material, with no animation or rigid bodies. The new test brief asked for 0.30 × 0.16 × 0 m within 0.001 m tolerance, plus a label legible from separate front and rear views.

Preparation proposed a dimensions check and two capture requests. The caller still had to review that interpretation. Running the measurements before approval was useful for diagnosis, but could not approve either the mapping or the scene.

Scroll sideways for more columns.

Part of the reportWhat ran or was requestedSaved result
General delivery policy27 selected rules; no task dimensions supplied to them.23 passed. Four had no applicable subjects: rigid bodies, colliders, joints and animated transforms.
Brief: panel dimensionsOne bounds comparison on /World/Panel.FAIL: width 0.24 m instead of 0.30 m, a 60 mm difference against 1 mm tolerance.
Brief: front legibilityDedicated front capture and visual review.UNKNOWN: no capture supplied.
Brief: rear legibilityA separate rear capture and visual review.UNKNOWN: no capture supplied.
Interpretation reviewCaller decision on the proposed requirement map.Pending. A correct-looking mapping is not its own approval.

The report counts subjects within each rule. Here, four rules had nothing to inspect; they provide no evidence of motion or physical behavior.

Recorded panel test. The delivered width is 240 millimetres against a 300 millimetre requirement and 1 millimetre tolerance. That comparison fails. Front and rear legibility remain unknown because the requested images are missing.
A diagram of the recorded requirement and result, not a scene render. Width can be measured from the USD. Legibility needs the caller's views. Neither question is answered by the 23 general-rule passes.

A separate repair control changed the panel to 0.30 m and updated its stored bounds. The check passed against the same approved scope, without another interpretation call. That is the behavior I want: repair the delivery against the requirement we agreed. Changing the policy or check implementations needs a new scope review.

The interpreter did not always organize the same brief identically. A repeat grouped it into two requirements instead of three, while retaining the dimensions and separate-view requests. A conflicting brief with widths of 0.30 and 0.40 m produced no width target and needed review. I would resolve the conflicting dimensions, then review and retain the proposed scope before using it to accept a delivery.

The application supplies the purpose and the evidence

An explanatory animation needs readable views and attached moving parts. Predicting motion under load adds dynamics and parameter evidence. The brief tells me which job I am accepting, and therefore which checks are worth running.

The caller passes a saved local USD bundle through check-3d-app or the Python library’s scene_acceptance.application.invoke; the 0.5 application protocol documents both interfaces. The caller owns the production loop. This adapter admits a restricted USD subset; assets using unsupported composition, such as payloads or instances, need another admission path.

Without a brief, it runs the caller’s selected general policy. With an already mapped brief, scripts can compare exact requirements without a model call. Raw text and images require an explicit preparation step with a configured interpreter. The model proposes checks, source references and capture expectations; the caller reviews them.

Application and harness responsibilities. The caller supplies a saved USD scene and optional brief. The harness can prepare proposed checks and requested views. The caller reviews the scope and renders the evidence. Scripts measure the scene and an optional judge reviews eligible evidence. The report returns findings and gaps; the caller decides whether to repair, provide evidence or accept. A repair is bound to the same approved scope.
Application flow in release 0.5, October 3. The intended use and brief determine which requirements need evidence; selected upstream rules and task-specific packs supply the checks. Rendering and its GPU belong to the caller. Scope approval and outcome acceptance are separate caller decisions.

In the press-shop demo, I asked for a realistic factory and later challenged the claim that the delivered video was photorealistic. Comparing decoded video frames with the original RTX images showed lost surface detail in the 1080p recording. The views were close but not exact pose matches, so this was a qualitative review rather than a pixel-aligned comparison. A separate, sharper RTX tooling still retained simplified machinery. Improving the recording would leave those scene limitations in place.

Original RTX tooling close-up from the earlier press-shop demo: broad metal surfaces, simple guide columns and feeder joints around the press opening.
Original RTX tooling still from the earlier 0.3 press-shop delivery, reviewed September 21. This is a direct scene render, not a frame from the compressed movie or the later improved scene. It helped inspect the machinery and surface detail; it does not measure recording loss. Capture provenance and review findings.

I want that gap visible in the acceptance report. It needs the actual frames and a clear visual requirement. The press-shop finding came from an assistant’s review after my challenge; whether the current optional judge would catch it remains untested.

Rendering stays with the application that created or opened the scene. The harness requests views, target parts, times and capture constraints. The caller returns the images with a receipt tied to the scene revision. A repaired scene needs fresh evidence for that revision.

Relabeling one panel image cannot satisfy both front and rear capture requests. Distinct images and matching metadata still leave a review question: do those views show the requested surfaces?

Two packaged skills guide the caller’s agent through evidence handoff and review customization. New numerical checks arrive as packs, with their parameters and coverage limits. The public polygon-budget example shows that extension. Installed packs run trusted Python.

The interpreter and the judge have different jobs

The interpreter proposes what to check. The judge makes a qualitative assessment of the candidate against that scope and the supplied evidence. By default, it does not see the script findings; the caller can choose to include them.

Scroll sideways for more columns.

QuestionScriptOptional LLM/VLM judgeInput still needed
Does the delivery meet selected USD and file rules?Reads composition, dependencies and selected mesh properties.Cannot replace these measurements.Application policy; no task brief required.
Are dimensions, placement or counts correct?Compares named objects with targets or a baseline.Can flag apparent layout concerns in suitable views.Targets, reference points and tolerances from the brief or contract. Images are not precise metrology.
Is timing or motion correct?Checks duration, sampled positions and named connection distances.Can flag an apparent separation or sequence problem in timed views.Required times and relationships. Neither proves unobserved intervals.
Does the texture meet the task?Resolves bindings, decodes images and compares selected source pixels.Reviews visible appearance or legibility in textured renders.Reference/comparison policy for pixels; suitable views for appearance.
Does the supported incline model meet its displacement limit?A CPU worker measures the fixed ramp-and-block model’s trace.Current judge has no physical-validation adapter; returns unknown.Supported model, task limit and parameter evidence.
Does it satisfy an unstated preference?No target to compare.Can raise a question, but cannot establish the owner’s unexpressed intent.Clarification or an explicitly chosen default.

The calling application selects checks, judge or both; checks are the default. Judge-only leaves numerical content checks unassessed. In both mode, the report keeps measured PASS/FAIL separate from advisory consistent/concern/unknown. A model opinion cannot override a required measurement failure.

The simulation worker covers one fixed ramp and rigid block with bounded parameter variations. For a different mechanism, I would need a suitable adapter and evidence that its model supports the intended use.

Default judge questions are selected from the scene inventory during preparation and remain fixed during repair. If a repair introduces materials or animation missing from that scope, the harness reports the coverage change and holds acceptance in judge or combined mode until a new scope is prepared. Checks-only results still disclose the omitted visual review.

In the panel’s no-image evaluation, the visual questions returned unknown without a judge call. A failed capture or model call also leaves completed script findings intact. I would rather have the width failure and an explicit request for views than lose a useful measurement because another component could not run.

What putting the harness to work taught me

I asked for fresh scenes at four brief depths: simple, light, medium and detailed. The study used three synthetic families, a gantry inspection cell, a rotary indexing bench and a pump/valve skid, with two attempts per condition. All 24 deliveries were readable and contained 62–70 geometry prims.

The requested model was gpt-6-astra with high reasoning and tools disabled. It returned USDA and a small palette raster that an adapter expanded into a PNG. Even the simple prompt supplied output, naming and coordinate conventions. No delivery was repaired, retried or given assessment feedback; the resolved model snapshot was unavailable. This was a constrained creation exercise, not an assessment of a complete tool-using Astra workflow.

Each unchanged delivery was checked against general policy, its supplied brief and a fuller target. I wanted to separate missed requirements from differences the brief had never forbidden.

Scroll sideways for more columns.

What happenedHow it was foundWhat I would change or keep
General policy rejected 21 of 24 deliveries.Missing stored mesh extents or required normals: 110 failed rule evaluations across 73 distinct scene/prim pairs.Run applicable delivery checks without waiting for a detailed brief. These are policy violations; some affected scenes still render.
All supplied geometry and motion rules passed; 5 of 6 detailed image rules failed.Exact source-image comparison found 768–4,096 different pixels out of 4,096.Keep checking supplied requirements after generation. Passing dimensions did not predict correct image content.
5 of 6 medium deliveries met their endpoints but missed the later full-path target.Intermediate positions were compared with a path withheld during creation.Supply the path if it matters. This result exposes ambiguity, not ignored instructions.
All 6 detailed paths passed at both 21 and 201 checkpoints.A denser scan found no additional failure in this cohort.Do not claim this study found a hidden between-sample defect. The constructed motion controls establish that risk separately.
Four simple/light skid clips did not reach later target times.The position checks returned unknown outside the authored interval.Ask for the missing time coverage; do not invent a position or count it as measured.
The visual judge established no additional scene defect.Follow-up inspection did not confirm two concerns about the same gantry cable; one motion concern matched an existing script failure.Keep concerns advisory and investigate them before counting defects. The value of the extra review cost is still unmeasured.

The general profile selected 27 rules. Three physics rules had no applicable subjects in every scene; no simulation ran in this creation study. The fuller target contained ten overlapping rules, so its pass count is not a percentage of total scene quality.

The judge study reused those 24 deliveries. Three repeats and two context probes brought it to 29 completed calls. The 24 primary reviews returned 51 consistent, 19 concern and 362 unknown items. Sixteen concerns involved occlusion or readability. The schematic views omitted textures, lighting and material transparency, leaving many questions unanswerable. Those unknowns do not establish either good or bad judge accuracy.

The scripts caught useful, specific failures. I still need task-appropriate renders and independent review of the judge’s concerns to decide whether its extra cost is worthwhile.

Read a failure as a request for a particular next step

The report retains the scene revision, inventory, brief source, selected checks, subject counts and unresolved requirements.

A failed dimension needs a repair. Missing front and rear views need captures. A disputed interpretation needs scope review. A crashed evaluator needs investigation. Sending all four back as “try generating the scene again” would waste attempts without necessarily resolving anything.

Dropping a check during repair must leave it unassessed. It cannot turn a previous failure into evidence that the problem was fixed.

Install it and run the tray example

The original v0.3 release provides a smaller reproducible test of the same acceptance idea. A translation-only task requires the tray to reach a destination while preserving its mesh. In the failing fixture, one interior floor point has moved 40 mm. The requested center and outer dimensions still match; six checks pass and shape preservation fails.

Two shaded views of the saved tray, at the same scale. Their outer dimensions match. A circled interior corner is unchanged in the good version and raised 40 millimetres in the bad version. An enlarged floor-edge profile reveals the difference. Shape preservation passes for the original and fails for the changed tray.
Both versions have the requested position and outer size. Comparing the saved mesh with the original catches the changed interior point. These views come from the USD fixtures; only the profile height is enlarged. The defect was deliberately constructed to test the evaluator.

That shape rule compares connectivity and corresponding points after removing translation. It suits a task that preserves the original representation. A valid remeshing or reindexing could fail it, so I would not use the same rule for an edit that permits those changes.

Download and extract the v0.3.0 source and evidence release. The pinned replay was tested with Python 3.12 on macOS arm64. From the extracted folder:

python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements-test.lock
python -m pip install --no-deps --no-build-isolation .

check-3d \
  --bundle-root evaluation/mesh-v1/fixtures/interior_shape_changed_same_bounds \
  --contract contract.json --candidate scene.usda \
  --baseline baseline.usda --out /tmp/tray-check-01

The expected result is REJECT, with exit code 2. Open /tmp/tray-check-01/report.html or inspect result.json. Choose a new output directory on each run; the CLI refuses to overwrite one or put it inside the input bundle. Relative scene and contract paths are resolved within --bundle-root.

The public replay uses OpenUSD 25.11 for scene reading and measurements, JSON Schema 4.25.1 for contracts and optional NVIDIA USD Validation 1.20.0 for selected upstream rules. The material, motion and physics articles each retain their versioned replay. The pack API carries the extension reference.

Other approaches also connect validation to intended use. SimReady Foundation includes profile validation and runtime testing. An adopter could build an acceptance workflow around those tools or an existing studio pipeline. This project tests one way to retain the job’s requirements, evidence and unresolved questions through delivery and repair.

The Astra crank-slider needed separate measurements of its linkage. Passing the static tray pack would say nothing about its moving attachments.

When I would take on the extra machinery

A small script tied the early harness on all 18 shared cube cases. For one stable conversion or edit, I would start there. Packs and a common report become useful when different jobs need different requirements, and the caller needs to distinguish a failed measurement from missing evidence without learning each evaluator’s output format.

The 0.5 package passed 668 software tests and the full reproduction command from a fresh archive extraction. Production savings remain unmeasured. The prepared producer trial still needs an independently supplied brief and external assessment, comparing time, compute and human review per accepted result against a competent script-and-review workflow, including unsuccessful attempts.

I intend to deepen the packs as my work needs them, using specialized evaluators where they fit. For now, I want every acceptance decision to show what was checked, against whose requirement, and what still needs evidence.

Disclaimer: The views and opinions expressed in this account are those of my own and do not represent those of my employer, NVIDIA.

← All projects