AI news ·
GPT-6 Astra hits 53% accuracy on indoor 3D scene coding but can't tell if its own work is correct
GPT-6 Astra scored only 53.4% accuracy reconstructing 3D indoor scenes from a single photo. The agents also failed to judge whether their own output improved, making human verification essential.

A research team from the University of Maryland and AWS developed a benchmark that measures how well AI coding agents can reconstruct 3D scenes from a single photograph - and the best model tested, GPT-6 Astra, reached only 53.4 percent accuracy on indoor scenes. The results expose a critical weakness: the agents cannot reliably judge whether their own output has improved, a limitation that directly affects anyone in product development, architecture, or simulation who might hope to generate editable 3D assets from reference images.
The project, called LEGO-Anything, uses an "Image-to-Code" approach. A coding agent receives a single image and writes code for Blender, the open-source 3D software. Instead of generating the scene in one pass, the agent works iteratively - it writes code, runs it, examines the result, and revises until the scene matches the original. Because the output is an executable program, it captures objects, geometry, layout, and camera position explicitly. You can run, check, and modify the scene like any other piece of code.
How LEGO-Bench measures reconstruction quality
To measure agent performance, the team built LEGO-Bench, a collection of 208 images from 104 indoor and outdoor scenes using 443 registered assets. Real photographs do not provide precise 3D ground truth for comparison, and simple synthetic scenes look unrealistic. The researchers split the difference by rendering images from professionally built simulator scenes. The inputs look natural, while exact geometry, depth, and object assignments remain hidden and serve as the answer key for automated scoring. Scene complexity can be increased without changing lighting or camera settings.
The benchmark scores each scene on three axes. Validity checks whether a usable scene artifact was delivered at all. Reconstruction measures how accurate the visible geometry is. Appearance captures how closely the look matches the original by re-rendering the submitted scene and comparing it pixel by pixel against the reference image.
Agents deliver working scenes but cannot self-assess
All six tested GPT configurations delivered a working scene almost every time. Accuracy varied sharply. GPT-6 Astra hit 53.4 percent on indoor scenes and 39.6 percent on outdoor scenes. Weaker configurations scored around 15 percent. Outdoor scenes proved harder than interiors, and accuracy dropped as scene complexity increased. When the researchers expanded the models' reasoning budget, GPT-6 variants improved substantially - Astra's score on an office test subset jumped from 32.3 to 61.8 percent.
The team analyzed the agents' work steps and found three recurring problems: poor initial attempts, revisions that undid earlier progress, and unreliable self-assessment. That last finding is the most revealing. When models had to choose which of two versions better matched the original, their geometric judgments landed near or below chance level. An agent essentially cannot tell whether its own scene has gotten better.
The researchers concluded that refinement should rely on concrete measurements, not the agent's own judgment. They built LEGO-Plugin, an extension requiring no extra training, which anchors the starting scene in the reference image, swaps unreliable self-judgment for concrete measurements, and shields correct progress from regressive edits. The plugin improved all six models. Weaker agents saw the biggest gains, with boosts up to 62.7 percent. The already strong top model gained roughly two percentage points.
Reconstructed scenes fall short for downstream tasks
The team also tested whether reconstructed scenes could support standard vision tasks. Because each scene is an executable program, object detection, segmentation, and depth estimation can be pulled directly from it without additional training. The scenes produced usable but unremarkable results across all three tasks. Object detection fared best, with reconstructed scenes reaching roughly half the performance of the specialized model DINO. The gap to specialized models like SAM 3 and Depth Anything 3 was larger for segmentation and depth estimation.
The authors said executable scene programs from current coding agents show promise but are not accurate enough. A wide gap remains between a working result and a faithful reconstruction. GPT-6 Astra's commanding lead in LEGO-Bench aligns with other observations. AI researcher Yoav Artzi sees the model as a major leap in spatial understanding and suspects it was trained on large amounts of 3D data such as Blender scenes. 3D software makers are already preparing for these kinds of agents - Unity has released official plugins for Claude Code and Codex. Other approaches skip code entirely and reconstruct scenes directly inside the model, such as the Atlas world model from World Labs. Google DeepMind takes yet another route with GenCeption, using a video model for depth estimation and segmentation that matches the performance of specialized models.
Why this matters for product development and research teams
For professionals in product development, architecture, or simulation, the ability to generate an editable 3D scene from a single photo would cut hours of manual modeling work. LEGO-Bench shows that current agents can produce a working Blender file almost every time, but the geometry is often wrong - and the agent does not know it. If you integrate these tools into a workflow today, you still need a human to verify spatial accuracy. The finding that concrete measurements dramatically improve weaker models points to a practical path forward: pairing agents with deterministic checks rather than trusting their self-assessment. For researchers and educators building on these methods, the benchmark and plugin are publicly available and provide a reproducible standard for measuring progress in image-to-code reconstruction.