← Projects

A Readable Texture Can Still Miss the Reference

A JPEG passed with one-level interior error and a 123-level whole-image error; both measurements were right

In this article 6 sections

One of the delivered JPEG labels passed its pattern check even though its largest pixel-channel error was 123 levels. The passing result was correct under the chosen comparison: it measured the interiors of four colored quadrants and excluded strips around their boundaries. Within that region, the largest error was one level.

An earlier summary called one level the image’s maximum without consistently saying “interior.” Checking the saved JPEG exposed the difference. The file was readable, the configured comparison passed, and the broader description of that pass was wrong.

That is the question I want this follow-up to answer: once the material and file-delivery checks pass, which parts of the image does the application actually require to match?

Six readable labels, two comparison policies

The September 24 evaluation used six supplied panel scenes from three delivery attempts: one PNG and one JPEG per attempt. The brief specified a 256 × 256 label with red, green, blue and white quadrants. All six images decoded successfully. Separate checks covered the shader connections and UV coverage.

My evaluating assistant compared every PNG pixel with the specified pattern. For JPEG, the policy allowed a maximum channel error of 32 levels and a mean of at most 4 per channel in the quadrant interiors. It excluded rows and columns 120–135 around their boundaries. That policy accommodates differences near the color transitions, but deliberately leaves those strips outside the claim.

Scroll sideways for more columns.

Supplied labelComparison regionLargest channel errorResult
All three PNGsWhole image0Exact pattern match.
JPEGs, attempts 1 and 2Whole image / selected interiors1 / 1Pass under the JPEG policy.
JPEG, attempt 3Whole image123Recorded diagnostic; not the selected acceptance region.
Same JPEG, attempt 3Selected interiors1Pass under the JPEG policy.

Errors here are absolute differences in decoded 8-bit RGB channel values. A value of 123 in the whole-image diagnostic does not contradict the one-level interior result. They measure different sets of pixels.

The JPEG comparison excludes a cross-shaped strip around the quadrant boundaries. A profile of decoded channel error shows large deviations near a boundary: whole-image maximum 123, compared with a maximum of 1 in the selected interiors.
The comparison mask shades excluded rows and columns. The profile shows error across row 64 of the saved third JPEG; its peak is 84 levels. The whole-image maximum is 123, while the selected interiors reach only 1. This is a pixel diagnostic, not a render of the panel.

I would not reuse this mask for a label whose boundary carries information. A narrow stripe, fine text or registration mark could sit in the excluded region. For that task, the acceptance policy must cover those pixels and set a tolerance suited to their use. Increasing tolerance or cropping away an awkward region after seeing a failure would change the requirement.

The six submitted images passed the intended checks. To test the evaluator, the assistant also made deliberately faulty copies. A missing file failed delivery. An unreadable file passed binding and file-presence checks but failed decoding. A readable image with the wrong pattern passed decoding and failed the pattern comparison. Those copies test detection; they are not defects found in the original six submissions.

A detailed brief did not eliminate wrong pixels

A later creation study produced 24 synthetic scenes at four brief depths. Six detailed deliveries received an exact 64 × 64 source-image reference. Five differed from it by 768–4,096 pixels out of 4,096, although their supplied geometry and motion checks passed.

These were unchanged producer deliveries. The requested Astra model had tools disabled and returned a small palette raster that the adapter expanded into PNG. No producer received assessment feedback or a repair attempt. The core study table records the wider conditions and keeps this cohort separate from the six September panels.

Exact matching was appropriate for those deterministic patterns. It would be a poor default for every photographic or compressed texture. The brief needs to identify the reference and what agreement means: exact pixels, selected regions, a declared tolerance, or a visible property requiring a rendered view. The harness cannot choose among those meanings merely because a file is named label.png.

Source-image agreement also leaves the final surface open. UV placement, occlusion, scale, lighting and color management can change what the viewer sees. Release 0.5 can request a caller-rendered view for optional visual review. That review is advisory; it cannot erase a failed source-pixel requirement.

The implementation in release 0.5 makes the comparison reusable: image_pixels accepts included or excluded rectangles, a declared black-and-white mask, and maximum and per-channel mean error limits. Its report retains the selected pixel count, mask hash, tolerances and whole-image diagnostics. These are decoded RGB channel comparisons, with no color-space conversion; they do not establish color-managed appearance. The earlier companion below retains its original controls.

Brief interpretation has a different image budget. It now accepts ordinary RGB or RGBA PNG/JPEG references up to 8 MiB and 16 million pixels each, within a 32 MiB total. Exact pixel checks keep their smaller RGB-only limit of 1 MiB and 262,144 pixels for both images, which must have matching dimensions. A larger image can be usable as a visual brief while needing a different evaluator for numerical comparison.

Read the pixels before comparing them

The decoder controls found a more basic trap: recognizing the format is not the same as reading the full image. One damaged PNG had enough of a header for Pillow to identify it and report dimensions. Verification failed later.

Pillow’s Image.open() is lazy. The check therefore verifies the file, reopens it and calls load() before reporting successful decoding. Plain text renamed to pixel.png fails earlier, during identification. A valid JPEG named pixel.png passes this particular contract because it checks the detected format, not agreement between the format and extension.

The integrated decoder uses Pillow rather than a new image reader. It records the selected shader attribute, resolved file, byte digest, format, dimensions and any failure. That ties a comparison to the delivered texture instead of whichever file happens to occupy the path later.

A correctly decoded red image still passes this decoder even when the task wants four quadrants. The reference comparison has to remain a separate selected requirement.

A check’s budget is not the producer’s size limit

The controls give the decoder a deliberately small budget of 1,024 pixels. A readable 33 × 33 PNG exceeds it and returns UNKNOWN. If decoding is required, the harness returns INSUFFICIENT_EVIDENCE. It has not established a content defect; this evaluator declined to inspect the image at that budget.

Compare that with a brief saying “deliver no more than 1,024 pixels.” The same 33 × 33 image would violate a delivery requirement. The number is identical, but it belongs to a different decision.

Scroll sideways for more columns.

ConditionWhat the report should sayUseful next action
Supported image has damaged bytesDecoding failed.Replace or repair the image.
Readable image exceeds evaluator budgetNot assessed within this budget.Increase the budget or select a suitable evaluator.
Image exceeds a stated delivery-size requirementSize requirement failed.Deliver a compliant image.
Readable image violates the supplied patternReference comparison failed.Repair the specified content, using the same comparison policy.

The experimental decoder covers single-frame PNG and JPEG. A valid GIF, animated PNG or file beyond its byte budget stays outside that coverage. Requiring the check makes these unresolved cases visible; making it advisory preserves the finding without blocking the other required checks.

Reproduce the comparison

The decoding pack and reference comparisons extend the material-delivery work in Scene Acceptance 0.5 on GitHub. The project also preserves the earlier valid/unreadable texture probe.

The v0.1 companion contains the earlier experimental pack, the harness source it was tested with, the original frozen bundles and their reports, with private environment paths removed from the public copy. This snapshot is based on harness commit 33e5721; it is separate from both v0.3.0 and the current 0.5 release. The later six-panel evaluation and supplemental checks remain in the private evaluation/fresh-producer-v1/ workspace; they are not included in this companion.

With Python 3.12 and the v0.1 companion extracted, run these commands from its root:

python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -e '.[test]' -e ./followups
python evaluation/followups-v1/run_matrix.py /tmp/harness-replay-01

Choose a new output directory for each run. The runner refuses to overwrite previous results and verifies the frozen input hashes before and after evaluation. In summary.json, the texture-* entries compare the delivery verdict, integrated verdict and direct Pillow observation. Each case also has the complete acceptance report.

The texture matrix contains eleven constructed cases. It was created and run with coding-assistant help under my direction, using OpenUSD 25.11 and Pillow 12.3.0. The companion retains the initial registration failure as well as corrected runs. These controls establish the behavior exercised here; they are not a representative sample of agent-generated textures or an independent reliability evaluation.

Put the comparison region in the result

For the JPEG case, I would keep the mask and tolerances with the reference, and report the interior and whole-image diagnostics with their names intact. “Pattern passed” was too easy to read as a statement about pixels the check had excluded.

The correction needed no new model or renderer. It needed a more precise account of the comparison we had already run.

Disclaimer: The views and opinions expressed in this account are those of my own and do not represent those of my employer, NVIDIA.

← All projects