LEGO-Anything: AI Agents Build 3D Scenes They Can't Judge

LEGO-Anything: AI Agents Build 3D Scenes They Can't Judge

Coding agents can now turn a single photo into an editable 3D scene by writing Blender code. A new benchmark from researchers at the University of Maryland and AWS shows that the agents usually deliver something that runs. It also shows that they cannot reliably tell whether their own work is improving.

From image to program

The project is called LEGO-Anything, and its method is called "Image-to-Code." A coding agent receives one image and writes code for Blender, the popular 3D software. It does not produce the scene in a single shot. It writes code, executes it, inspects the render and revises, repeating the loop until the result resembles the input.

The output is a program, not a mesh or a picture. Objects, geometry, layout and camera position are all spelled out explicitly. That means the scene can be run, inspected, queried and edited like any other code.

A benchmark with a hidden answer key

To score the agents, the team built LEGO-Bench. It has 208 images taken from 104 indoor and outdoor scenes, built from 443 registered assets.

The design addresses a basic problem. Real photos come with no exact 3D ground truth, and simple synthetic scenes look artificial. LEGO-Bench renders its images from professionally built simulator scenes instead. The inputs look natural, while the precise geometry, depth and object assignments stay hidden and act as the answer key. The researchers can also make scenes more complex without changing lighting or camera settings.

Each submission is graded on three axes:

  • Validity: did the agent deliver a usable scene at all?
  • Reconstruction: how accurate is the visible geometry?
  • Appearance: how closely does a re-render of the scene match the reference image, pixel by pixel?

Working output, weak geometry

All six GPT configurations tested produced a working scene almost every time. Accuracy was a different story. The strongest model, GPT-6 Astra, reached 53.4 percent on indoor scenes and 39.6 percent on outdoor ones. Weaker configurations landed around 15 percent.

Accuracy fell as scenes got more complex, and outdoor scenes were harder than interiors. More reasoning budget helped the GPT-6 variants considerably. On an office test subset, Astra climbed from 32.3 to 61.8 percent.

The researchers then examined the agents' step-by-step work. The most common problems were poor first attempts, revisions that erased earlier progress and unreliable self-assessment.

The self-assessment result is the most striking. Asked to choose which of two versions better matched the original, the models made geometric judgments at or below chance level. In practice, an agent cannot tell whether its scene got better or worse. The authors conclude that refinement should be driven by concrete measurements, not by the model's own opinion.

A plugin that replaces self-judgment

That finding led to LEGO-Plugin, an add-on that requires no additional training. It anchors the initial scene in the reference image, substitutes concrete measurements for the agent's self-evaluation and protects correct progress from edits that would undo it. All six models improved. Weaker agents gained the most, with boosts of up to 62.7 percent, while the top model added only about two percentage points.

Not yet good enough for vision tasks

The team also checked whether the reconstructed scenes could feed standard vision tasks. Because each scene is code, object detection, segmentation and depth estimation can be extracted directly. Without extra training, results were usable but modest. Detection did best, reaching roughly half the performance of the specialized model DINO. The gap to SAM 3 and Depth Anything 3 on segmentation and depth was wider. The authors describe the approach as promising but not yet accurate enough.

Our Take

The core lesson echoes what we keep seeing in agent research: a model can be capable and still be a poor judge of its own output. LEGO-Plugin's gains came from moving evaluation outside the model, which suggests the same pattern may pay off in other agent workflows.

The field is also splitting into routes. Unity has released official plugins for Claude Code and Codex, and code-based tooling around agents like Claude Code keeps growing. Others skip code entirely, such as the Atlas world model from World Labs, while Google Deepmind's GenCeption uses a video model for depth and segmentation. It is worth watching whether editable code scenes can close the accuracy gap, or whether direct world models win on fidelity while code wins on control.