Building an Acceptance Harness for Agent-Generated 3D
Working backward from a brief to the checks and evidence a 3D delivery needs
When a 3D job is delivered, I want to know: is it ready for its intended use, and what evidence supports that decision? “The scene passed” was hiding too much. A file could open, satisfy the general delivery rules and still miss a dimension in the brief. A label could have the right source pixels while nobody had checked whether it was readable on the object.
I built this for my own 3D work, starting with the outcomes I needed and working backward to the checks that could support them. That brought geometry, materials, motion, rendered views and some simulation into one acceptance workflow. I spent time on the parts that mattered to those jobs, so the depth varies across them. A different use may need much more from one area and nothing from another.
OpenUSD validators and NVIDIA’s USD validation library already provide extensible ways to check assets. This harness uses selected rules from usd-validation-nvidia==1.20.0 and reports their findings alongside a job’s requirements. Where a task needs a specialist validator or simulator, I would use that through a suitable pack. Bringing the results together does not make each check as deep as a dedicated tool.
Version scope: Scene Acceptance 0.5 includes the interpreter, optional judge and application CLI described here. The studies used local revision 52e412f, tested October 3, 2026; their raw packets remain private and retain their original policy. Later 0.5 fixes have separate regression tests. The public v0.3 release supports the tray replay below; the motion article links the public timing fix.
I asked my coordinating assistant to implement the harness and compare no brief, increasingly detailed briefs and optional model review. The assistant authored the synthetic briefs and evaluated the results. These are exploratory studies, not an independent reliability benchmark.
One panel, a width failure and two missing views
The interpreter test used an existing panel: seven prims, including one mesh and one material, with no animation or rigid bodies. The new test brief asked for 0.30 × 0.16 × 0 m within 0.001 m tolerance, plus a label legible from separate front and rear views.
Preparation proposed a dimensions check and two capture requests. The caller still had to review that interpretation. Running the measurements before approval was useful for diagnosis, but could not approve either the mapping or the scene.
Scroll sideways for more columns.
| Part of the report | What ran or was requested | Saved result |
|---|---|---|
| General delivery policy | 27 selected rules; no task dimensions supplied to them. | 23 passed. Four had no applicable subjects: rigid bodies, colliders, joints and animated transforms. |
| Brief: panel dimensions | One bounds comparison on /World/Panel. | FAIL: width 0.24 m instead of 0.30 m, a 60 mm difference against 1 mm tolerance. |
| Brief: front legibility | Dedicated front capture and visual review. | UNKNOWN: no capture supplied. |
| Brief: rear legibility | A separate rear capture and visual review. | UNKNOWN: no capture supplied. |
| Interpretation review | Caller decision on the proposed requirement map. | Pending. A correct-looking mapping is not its own approval. |
The report counts subjects within each rule. Here, four rules had nothing to inspect; they provide no evidence of motion or physical behavior.
A separate repair control changed the panel to 0.30 m and updated its stored bounds. The check passed against the same approved scope, without another interpretation call. That is the behavior I want: repair the delivery against the requirement we agreed. Changing the policy or check implementations needs a new scope review.
The interpreter did not always organize the same brief identically. A repeat grouped it into two requirements instead of three, while retaining the dimensions and separate-view requests. A conflicting brief with widths of 0.30 and 0.40 m produced no width target and needed review. I would resolve the conflicting dimensions, then review and retain the proposed scope before using it to accept a delivery.
The application supplies the purpose and the evidence
An explanatory animation needs readable views and attached moving parts. Predicting motion under load adds dynamics and parameter evidence. The brief tells me which job I am accepting, and therefore which checks are worth running.
The caller passes a saved local USD bundle through check-3d-app or the Python library’s scene_acceptance.application.invoke; the 0.5 application protocol documents both interfaces. The caller owns the production loop. This adapter admits a restricted USD subset; assets using unsupported composition, such as payloads or instances, need another admission path.
Without a brief, it runs the caller’s selected general policy. With an already mapped brief, scripts can compare exact requirements without a model call. Raw text and images require an explicit preparation step with a configured interpreter. The model proposes checks, source references and capture expectations; the caller reviews them.
In the press-shop demo, I asked for a realistic factory and later challenged the claim that the delivered video was photorealistic. Comparing decoded video frames with the original RTX images showed lost surface detail in the 1080p recording. The views were close but not exact pose matches, so this was a qualitative review rather than a pixel-aligned comparison. A separate, sharper RTX tooling still retained simplified machinery. Improving the recording would leave those scene limitations in place.
I want that gap visible in the acceptance report. It needs the actual frames and a clear visual requirement. The press-shop finding came from an assistant’s review after my challenge; whether the current optional judge would catch it remains untested.
Rendering stays with the application that created or opened the scene. The harness requests views, target parts, times and capture constraints. The caller returns the images with a receipt tied to the scene revision. A repaired scene needs fresh evidence for that revision.
Relabeling one panel image cannot satisfy both front and rear capture requests. Distinct images and matching metadata still leave a review question: do those views show the requested surfaces?
Two packaged skills guide the caller’s agent through evidence handoff and review customization. New numerical checks arrive as packs, with their parameters and coverage limits. The public polygon-budget example shows that extension. Installed packs run trusted Python.
The interpreter and the judge have different jobs
The interpreter proposes what to check. The judge makes a qualitative assessment of the candidate against that scope and the supplied evidence. By default, it does not see the script findings; the caller can choose to include them.
Scroll sideways for more columns.
| Question | Script | Optional LLM/VLM judge | Input still needed |
|---|---|---|---|
| Does the delivery meet selected USD and file rules? | Reads composition, dependencies and selected mesh properties. | Cannot replace these measurements. | Application policy; no task brief required. |
| Are dimensions, placement or counts correct? | Compares named objects with targets or a baseline. | Can flag apparent layout concerns in suitable views. | Targets, reference points and tolerances from the brief or contract. Images are not precise metrology. |
| Is timing or motion correct? | Checks duration, sampled positions and named connection distances. | Can flag an apparent separation or sequence problem in timed views. | Required times and relationships. Neither proves unobserved intervals. |
| Does the texture meet the task? | Resolves bindings, decodes images and compares selected source pixels. | Reviews visible appearance or legibility in textured renders. | Reference/comparison policy for pixels; suitable views for appearance. |
| Does the supported incline model meet its displacement limit? | A CPU worker measures the fixed ramp-and-block model’s trace. | Current judge has no physical-validation adapter; returns unknown. | Supported model, task limit and parameter evidence. |
| Does it satisfy an unstated preference? | No target to compare. | Can raise a question, but cannot establish the owner’s unexpressed intent. | Clarification or an explicitly chosen default. |
The calling application selects checks, judge or both; checks are the default. Judge-only leaves numerical content checks unassessed. In both mode, the report keeps measured PASS/FAIL separate from advisory consistent/concern/unknown. A model opinion cannot override a required measurement failure.
The simulation worker covers one fixed ramp and rigid block with bounded parameter variations. For a different mechanism, I would need a suitable adapter and evidence that its model supports the intended use.
Default judge questions are selected from the scene inventory during preparation and remain fixed during repair. If a repair introduces materials or animation missing from that scope, the harness reports the coverage change and holds acceptance in judge or combined mode until a new scope is prepared. Checks-only results still disclose the omitted visual review.
In the panel’s no-image evaluation, the visual questions returned unknown without a judge call. A failed capture or model call also leaves completed script findings intact. I would rather have the width failure and an explicit request for views than lose a useful measurement because another component could not run.
What putting the harness to work taught me
I asked for fresh scenes at four brief depths: simple, light, medium and detailed. The study used three synthetic families, a gantry inspection cell, a rotary indexing bench and a pump/valve skid, with two attempts per condition. All 24 deliveries were readable and contained 62–70 geometry prims.
The requested model was gpt-6-astra with high reasoning and tools disabled. It returned USDA and a small palette raster that an adapter expanded into a PNG. Even the simple prompt supplied output, naming and coordinate conventions. No delivery was repaired, retried or given assessment feedback; the resolved model snapshot was unavailable. This was a constrained creation exercise, not an assessment of a complete tool-using Astra workflow.
Each unchanged delivery was checked against general policy, its supplied brief and a fuller target. I wanted to separate missed requirements from differences the brief had never forbidden.
Scroll sideways for more columns.
| What happened | How it was found | What I would change or keep |
|---|---|---|
| General policy rejected 21 of 24 deliveries. | Missing stored mesh extents or required normals: 110 failed rule evaluations across 73 distinct scene/prim pairs. | Run applicable delivery checks without waiting for a detailed brief. These are policy violations; some affected scenes still render. |
| All supplied geometry and motion rules passed; 5 of 6 detailed image rules failed. | Exact source-image comparison found 768–4,096 different pixels out of 4,096. | Keep checking supplied requirements after generation. Passing dimensions did not predict correct image content. |
| 5 of 6 medium deliveries met their endpoints but missed the later full-path target. | Intermediate positions were compared with a path withheld during creation. | Supply the path if it matters. This result exposes ambiguity, not ignored instructions. |
| All 6 detailed paths passed at both 21 and 201 checkpoints. | A denser scan found no additional failure in this cohort. | Do not claim this study found a hidden between-sample defect. The constructed motion controls establish that risk separately. |
| Four simple/light skid clips did not reach later target times. | The position checks returned unknown outside the authored interval. | Ask for the missing time coverage; do not invent a position or count it as measured. |
| The visual judge established no additional scene defect. | Follow-up inspection did not confirm two concerns about the same gantry cable; one motion concern matched an existing script failure. | Keep concerns advisory and investigate them before counting defects. The value of the extra review cost is still unmeasured. |
The general profile selected 27 rules. Three physics rules had no applicable subjects in every scene; no simulation ran in this creation study. The fuller target contained ten overlapping rules, so its pass count is not a percentage of total scene quality.
The judge study reused those 24 deliveries. Three repeats and two context probes brought it to 29 completed calls. The 24 primary reviews returned 51 consistent, 19 concern and 362 unknown items. Sixteen concerns involved occlusion or readability. The schematic views omitted textures, lighting and material transparency, leaving many questions unanswerable. Those unknowns do not establish either good or bad judge accuracy.
The scripts caught useful, specific failures. I still need task-appropriate renders and independent review of the judge’s concerns to decide whether its extra cost is worthwhile.
Read a failure as a request for a particular next step
The report retains the scene revision, inventory, brief source, selected checks, subject counts and unresolved requirements.
A failed dimension needs a repair. Missing front and rear views need captures. A disputed interpretation needs scope review. A crashed evaluator needs investigation. Sending all four back as “try generating the scene again” would waste attempts without necessarily resolving anything.
Dropping a check during repair must leave it unassessed. It cannot turn a previous failure into evidence that the problem was fixed.
Install it and run the tray example
The original v0.3 release provides a smaller reproducible test of the same acceptance idea. A translation-only task requires the tray to reach a destination while preserving its mesh. In the failing fixture, one interior floor point has moved 40 mm. The requested center and outer dimensions still match; six checks pass and shape preservation fails.
That shape rule compares connectivity and corresponding points after removing translation. It suits a task that preserves the original representation. A valid remeshing or reindexing could fail it, so I would not use the same rule for an edit that permits those changes.
Download and extract the v0.3.0 source and evidence release. The pinned replay was tested with Python 3.12 on macOS arm64. From the extracted folder:
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements-test.lock
python -m pip install --no-deps --no-build-isolation .
check-3d \
--bundle-root evaluation/mesh-v1/fixtures/interior_shape_changed_same_bounds \
--contract contract.json --candidate scene.usda \
--baseline baseline.usda --out /tmp/tray-check-01
The expected result is REJECT, with exit code 2. Open /tmp/tray-check-01/report.html or inspect result.json. Choose a new output directory on each run; the CLI refuses to overwrite one or put it inside the input bundle. Relative scene and contract paths are resolved within --bundle-root.
The public replay uses OpenUSD 25.11 for scene reading and measurements, JSON Schema 4.25.1 for contracts and optional NVIDIA USD Validation 1.20.0 for selected upstream rules. The material, motion and physics articles each retain their versioned replay. The pack API carries the extension reference.
Other approaches also connect validation to intended use. SimReady Foundation includes profile validation and runtime testing. An adopter could build an acceptance workflow around those tools or an existing studio pipeline. This project tests one way to retain the job’s requirements, evidence and unresolved questions through delivery and repair.
The Astra crank-slider needed separate measurements of its linkage. Passing the static tray pack would say nothing about its moving attachments.
When I would take on the extra machinery
A small script tied the early harness on all 18 shared cube cases. For one stable conversion or edit, I would start there. Packs and a common report become useful when different jobs need different requirements, and the caller needs to distinguish a failed measurement from missing evidence without learning each evaluator’s output format.
The 0.5 package passed 668 software tests and the full reproduction command from a fresh archive extraction. Production savings remain unmeasured. The prepared producer trial still needs an independently supplied brief and external assessment, comparing time, compute and human review per accepted result against a competent script-and-review workflow, including unsuccessful attempts.
I intend to deepen the packs as my work needs them, using specialized evaluators where they fit. For now, I want every acceptance decision to show what was checked, against whose requirement, and what still needs evidence.
Disclaimer: The views and opinions expressed in this account are those of my own and do not represent those of my employer, NVIDIA.