Oct 3, 2026
Nano Banana Pro prompted by THE DECODER
Coding agents can build an editable 3D scene from a single photo by writing and refining code step by step. The benchmark that comes with this work reveals where the process still breaks down: self-assessment and geometric accuracy.
A 3D reconstruction from a single photo is most useful when it exists as an executable program you can inspect, edit, and query. That's the premise behind LEGO-Anything, a project from researchers at the University of Maryland and AWS.
LEGO-Anything has a coding agent write an executable Blender program from a single photo. The resulting scene can be edited and queried for image analysis tasks. | Image: Li et al.
The approach is called "Image-to-Code." A coding agent receives a single image and writes code for Blender, the widely used 3D software. Rather than generating the scene in one pass, the agent works iteratively: it writes code, runs it, looks at the result, and revises until the scene matches the original.
Because the output is a program, it captures objects, geometry, layout, and camera position explicitly. You can run, check, and modify the scene like any other piece of code.
Simulator scenes provide the exact ground truth
To measure how well agents perform, the team introduces LEGO-Bench. It contains 208 images from 104 indoor and outdoor scenes and uses 443 registered assets.
Real photos don't provide a precise 3D ground truth to compare against, according to the researchers. Simple synthetic scenes look unrealistic. So LEGO-Bench splits the difference by rendering its images from professionally built simulator scenes. The inputs look natural, while the exact geometry, depth, and object assignments stay hidden and serve as the answer key for automated scoring. Scene complexity can also be ramped up without changing lighting or camera settings.
LEGO-Bench scores scenes on validity, geometric accuracy, and visual similarity to the original. | Image: Li et al.
The benchmark scores each scene on three axes: validity checks whether a usable scene artifact was delivered at all. Reconstruction measures how accurate the visible geometry is. Appearance captures how closely the look matches the original by re-rendering the submitted scene and comparing it pixel by pixel against the reference image.
Agents deliver usable artifacts but struggle with geometry
All six tested GPT configurations delivered a working scene almost every time. Accuracy varied wildly, though. GPT-6 Astra, the best tested model, hit 53.4 percent on indoor scenes and 39.6 percent on outdoor scenes. Weaker configurations scored around 15 percent.
GPT-6 Astra comes closest to the reference images. Older models frequently miss camera angles, lighting, or entire objects. | Image: Li et al.
The more complex a scene, the more accuracy drops, and outdoor scenes are harder than interiors. When the researchers increased the models' reasoning budget, the GPT-6 variants improved a lot. Astra's score on an office test subset jumped from 32.3 to 61.8 percent.
To understand why, the researchers analyzed the agents' work steps. They found poor initial attempts, revisions that undid earlier progress, and unreliable self-assessment as the most common issues.
Even GPT-6 Astra wrecks its own scene late in the process, dropping from 33.9 to 4.4 percent. | Image: Li et al.
That last point is the most telling. When models had to pick which of two versions better matched the original, their geometric judgments landed near or below chance level. An agent basically can't tell whether its own scene has gotten better. The researchers conclude that refinement should rely on concrete measurements, not the agent's own judgment.
When judging geometry, the models perform near chance level, even when evaluating their own scenes. | Image: Li et al.
That insight led the authors to build LEGO-Plugin, an extension that needs no extra training. It anchors the starting scene in the reference image, swaps the unreliable self-judgment for concrete measurements, and shields correct progress from regressive edits. The plugin improved all six models. Weaker agents saw the biggest gains, with boosts up to 62.7 percent. The already strong top model gained only about two percentage points.
The plugin improves all tested models. Weaker agents benefit the most. | Image: Li et al.
Reconstructed scenes aren't accurate enough yet
Finally, the team tested whether reconstructed scenes could serve as a basis for standard vision tasks. Because each scene is an executable program, object detection, segmentation, and depth estimation can be pulled directly from it.
Without any extra training, the scenes produced usable but unremarkable results across all three tasks. Object detection fared best, with the reconstructed scenes reaching roughly half the performance of the specialized model DINO. For segmentation and depth estimation, the gap to specialized models like SAM 3 and Depth Anything 3 was larger. The authors say executable scene programs from current coding agents show promise but aren't accurate enough. A big gap remains between a working result and a faithful reconstruction.
GPT-6 Astra's commanding lead in LEGO-Bench lines up with other observations. AI researcher Yoav Artzi sees the model as a major leap in spatial understanding and suspects it was trained on large amounts of 3D data like Blender scenes. 3D software makers are already gearing up for these kinds of agents. Unity has released official plugins for Claude Code and Codex. Other approaches skip code entirely and reconstruct scenes directly inside the model, like the Atlas world model from World Labs. Google Deepmind takes yet another route with GenCeption, using a video model for depth estimation and segmentation that matches the performance of specialized models.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Read on for the full picture.
Subscribe for hype-free coverage.
-
Full access to every article on THE DECODER
-
No ads
-
Join the comments and community discussions
-
A weekly AI news recap via mail
-
6x/year: "AI Radar" — deep dives on the AI topics that matter most
-
Daily AI news, always up to date
-
Our full ten-year archive
-
Covered by a team with 10+ years in AI