← Writing
Delegating 3D Production Part 3 of 3

What Does a Hallucination Look Like in 3D?

A correct coordinate, an unsupported claim, and five probes that found no hallucination

In this article 7 sections

When an agent creates 3D content, inventing an object may be part of the brief. I would call it a spatial hallucination when it presents an object, relationship or physical property as established by evidence that does not support it. The claim may appear in a report or plan even when the scene looks plausible.

The pallet-change review checked whether a saved edit fulfilled the request. Here I examine what can be claimed about the result and what should remain unknown. I had asked whether hallucination was a useful concept in 3D, then asked for an actual example we could record. None of the five local text-only Astra trials produced one.

“Friction and mass are unspecified.” That was Astra’s answer to a question about a conveyor deck shown in a short USD scene description. The prompt asked whether it would support a carton, and what its mass and friction were.

The excerpt supplied a visible cube, dimensions and a color. It supplied no collision schema or authored mass or friction. Astra said contact support was not established and left the missing properties unresolved. It also observed that a static collider need not have an assigned dynamic mass. That was a useful answer to an incomplete description.

Projection derived from trial 4's supplied USD cube: a deck 2 by 0.7 by 0.1 metres, with its top at Z equals 0.90 metres. That coordinate is supported. An illustrative claim that the excerpt configures a collider is contradicted because no PhysicsCollisionAPI is applied. Whether a carton remains supported is unknown from this evidence; friction and mass are unspecified. These are claims to assess, not three claims Astra made.
Derived from the USD text in trial 4, with diagram shading and dimension labels added. The model received text only. The collider claim is an illustrative claim to assess; Astra did not make it. Its answer correctly left contact support unresolved.

What I would accept also depends on the job. A plausible animation can complete a visual brief while leaving the evidence needed for a physical prediction unresolved.

What happened when we tried Astra

The coding agent prepared five text prompts and expected checks before submitting them through Codex on 19 September 2026, requesting gpt-6-astra with high reasoning. Each used a fresh session and no tools. The questions covered geometry, transforms, missing physics and an old validation report.

All completed responses are summarized below. The companion evidence bundle contains every exact prompt and full answer, selected run metadata and the deterministic calculations used to check them.

Scroll sideways for more columns.

TrialQuestion testedAstra’s responseVerification
1Does a rotated pallet enter the aisle?No; minimum Y = 1.24699 m, leaving 0.04699 m clearance.Transformed corners agree.
2Do two diagonal bars intersect?No; their closest surfaces are 0.012132 m apart.Oriented-axis projections confirm the gap despite overlapping axis-aligned bounds.
3Does equipment in a rotated parent assembly enter the aisle?Yes; minimum world Y = 1.10 m, giving 0.10 m intrusion.Transforming all eight corners agrees.
4Does a visible deck establish support, mass and friction?Contact support is not established; mass and friction are unspecified.The supplied USD has no collision schema or authored physical properties.
5Does an earlier passing report justify accepting a moved pallet?No; the current placement overlaps the aisle by 0.30 m.Current coordinates confirm the violation; the report covers the previous revision.

That is four supported geometric or revision judgments and one appropriate statement of missing information. There was one response per question, with no completed-response reruns. A sixth planned image-based case has not run and is excluded from the results.

The assistant designed the prompts and assessed the answers through geometric calculations and inspection of the supplied USD. The bundle lets a reader repeat those checks; it does not rerun the stochastic model. No immutable model snapshot was returned. Five successes on different supplied-text questions cannot establish reliability across photographs, long authoring sessions or another model.

I would be cautious about how much weight to put on the fourth answer. Its prompt explicitly restricted the model to the supplied excerpt and ruled out other layers or simulation settings, which gave it a favorable setup for recognizing missing information. We have not shown that the same restraint survives a long authoring session, conflicting context or an ambiguous image.

Supported, contradicted and not established

I would assess claims about the saved scene using three outcomes.

The second and third rows below are claims to assess. Astra rejected the stale-report claim and left the physics question unresolved; it did not assert either as fact.

Scroll sideways for more columns.

Claim under reviewEvidence availableJudgment and next action
“This rotated pallet clears the authored aisle.”Trial 1’s supplied geometry and verified transformed cornersSupported for that geometry. Keep the calculation and its inputs with the result.
“The moved pallet is clear because its old report passed.”Trial 5’s current coordinates give 0.30 m intrusionContradicted. Reject that clearance claim and rerun checks for the current revision.
“This visible deck will support the carton with a particular friction value.”Trial 4 supplies geometry but no configured collision or friction evidenceNot established. Obtain the missing configuration and relevant test before making that claim.

I would handle the second and third rows differently. The current coordinates are enough to reject the clearance claim, while the missing physical configuration leaves a question to resolve before making a claim about support. Saying “it definitely cannot support a carton” would also go beyond the evidence if the question concerns a larger system whose configuration we have not seen.

An evaluator that labels everything “unknown” would avoid confident errors but answer none of these questions. I would test supported claims, contradictions and deliberately incomplete inputs separately. These five responses do not establish a reliable classifier.

The same scene, different acceptance criteria

For the conveyor in the fourth trial, the missing properties matter differently depending on what I ask it to do. An animation explaining a proposed packaging process may be complete with the requested objects, sequence and convincing carton motion. Using that same scene to predict whether the carton slips requires evidence about the contact. The animation has not supplied it.

An unknown therefore does not make every use of the scene unacceptable. It leaves the affected claim unresolved. A visualization used to approve actual plant clearance still needs dimensional evidence even if nobody runs a physics simulation. The intended decision determines which missing information blocks acceptance.

For the conveyor, I would retain an assumed friction value as an assumption and decide whether the intended use can tolerate that uncertainty. Calling it measured would be unsupported in every use case.

A reported failure, with a different model

There is an actual object-grounding failure in Figure 5 of the 3D-VCD preprint. In a HEAL task involving removing lint from clothing, Qwen-14B-Instruct included microwave.n.01_1 in its predicted symbolic goals. The authors identify that microwave as absent from the scene. Their modified decoding method omitted it.

The evaluation introduced a distracting object through the task description. The failure was carrying that distraction into a goal about the environment. It was a well-formed symbolic output containing an unsupported object reference, rather than a visibly broken mesh.

This is the researchers’ reported result from another model and an adversarial benchmark. I have not reproduced it, and it tells us neither Astra’s failure rate nor how often a coding agent would make a similar mistake during scene authoring.

The example gives me a concrete question for a planner: does the referenced object exist in the available scene evidence? That evidence itself has a boundary. An object absent from a complete task inventory is different from an object that was simply outside one camera’s view. The latter can remain unknown.

A correct number can acquire an unsupported meaning

The packaging-cell exercise contains an authored 1.20 metre aisle requirement. Its alternative placement has a computed pallet-edge coordinate of Y = 1.25 metres. Neither number came from surveying a plant.

Three claims with different evidence: 1.20 metres is the authored aisle requirement; 1.25 metres is the pallet edge computed from the saved scene; actual plant clearance is unknown because no survey was performed.
The numbers support statements about this authored layout. Actual facility clearance remains unknown.

If a generated report calls the result a measured plant clearance, the geometry need not change for the claim to become unsupported. This is a hypothetical reporting error, not something Astra did in our tests. The extra meaning came from the report.

I would therefore retain the basis of each consequential property: supplied, measured, computed or assumed, with a source and a scene revision. A computed number may depend on assumed dimensions. Labeling it “computed” does not remove those assumptions or make it a field measurement.

Diagnose the failure before choosing the repair

The earlier table asks whether a claim is supported. This one asks what went wrong and what to do next. These are illustrative scenarios for diagnosis, not additional observed model failures.

Scroll sideways for more columns.

ExampleDiagnosisNext step
The pallet moves to a feasible position that the brief did not request.The completion check omitted the target. This happened in the constructed pallet review.Add the destination requirement; resolve any conflict with the aisle constraint.
A report calls the computed pallet edge a measured plant clearance.Unsupported claim about the source of the number. This is a hypothetical reporting error.Correct the provenance; obtain survey evidence if the use needs it.
A plan names the absent microwave in the published example above.Unsupported object reference under that benchmark’s complete inventory.Check the plan’s references against the task inventory.
The deck’s friction is unspecified and the answer says so.Missing evidence, correctly preserved in trial 4.Decide whether the intended use needs a contact model and measured parameters.

These labels are not mutually exclusive. A construction error can be followed by an unsupported claim that the result was verified. Missing information is not automatically the user’s fault; the agent may need to request it or preserve the uncertainty. An object outside a partial camera view is also not known to be absent.

I would compare the brief, available inputs, saved artifact and resulting claims before assigning a cause. An upstream perception error or faulty tool can supply a bad premise. If the evidence cannot distinguish those causes, I would leave the diagnosis unresolved and state what is needed to investigate it.

Keep the claim with its evidence

For a consequential property, I would keep the value, its source, how it was obtained and the scene revision together. A correct coordinate does not justify a field-measurement claim, and a missing camera view does not prove an object is absent.

The related Scene Acceptance project on GitHub implements contract-selected checks and reports their findings and unresolved evidence. It supplies the acceptance component of this discussion; the five Astra question-and-answer probes remain in the companion download. Release 0.5 includes optional model review, while the later study packets remain private.

The later harness study reinforced the need to investigate model opinions: two apparent cable concerns were not confirmed as source breaks when the saved geometry was inspected. I would retain the concern and that follow-up together. Calling either one a detected hallucination would claim more than the test established.

Disclaimer: The views and opinions expressed in this account are those of my own and do not represent those of my employer, NVIDIA.

← All writing