Choose a prompt, reference images, a blueprint set or an admitted capture. Select the target cell and build recipe, then keep the source, preflight and run evidence in the same spatial project.
The current release handles source admission, recipe selection, preflight and run inspection. Full prompt-to-world provider execution remains an integration target.
The first capture experiment began from a generated front elevation. It predates the measured structure and exact-camera FLUX studies documented in the Case Study.
A traced plan and style brief produced this image. The evaluation loop compares candidate motion, solves cameras, trains registered views and scores held-out views. Its narrow front coverage defined the need for exact multi-angle references.
The source image shows 2 stories, 4 porch columns, 3 balcony bays, a five-sided brick bay tower on the right and a single-story glazed sun-room on the left. The asymmetry distinguishes the intended orientation from mirrored or structurally changed frames.
All seven ran from the identical generated still with the same prompt and negative prompt, then faced a local screening proxy: feature continuity, camera motion, source retention, cut safety. Thirteen runs across two scenes, $8.44 actual against a $9.00 budget.
The proxy ranked Veo first at 90.9 and Kling fifth at 85.3. Structure from motion reversed that order: Kling reconstructed and Veo failed. The proxy filters obvious visual failures. Camera recovery determines whether a clip contains usable multi-view geometry.
Screening formula: 35% feature continuity, 25% camera motion, 25% source retention and 15% cut safety over SIFT correspondences with RANSAC fundamental geometry. Persistent 2D texture can score without coherent multi-view structure, so every passing clip also receives a camera solve.
Each clip goes to structure from motion under pinned intrinsics, and the score counts how many of 210 image pairs return a geometrically calibrated relationship. A clip with no stable camera solution stops before any paid generation or reconstruction training.
Camera recovery and visual quality are separate review gates. The high-scoring a150 guide failed on roofline and foliage artifacts. The pb100 guide failed local camera recovery and stopped before generation.
The first guide that produced a solvable generation at all.
The middle arc of the r7 family, selected for the next generation.
Highest guide solve score; rejected for roofline bending, foliage smear and five watermark pairs.
The guide calibrated 190 of 210 pairs; its generated clip calibrated 206.
Pure pull-back motion supplied insufficient lateral parallax for mapper initialization, so the candidate stopped after the local solve.
Six candidates were rendered and solved locally at no cost. Five passed at 21/21 registered, 0.73 to 0.82 px, zero watermark-configuration pairs. One failed outright. Paid generation proceeded for the five passing guides.
Visual review passed, but every pairwise solve collapsed into one degenerate camera configuration.
The first generated clip that reconstructed at all.
Sharpest within its solved arc; unsupported views collapse outside that range.
Strongest parallax in the project, and the first view past the corner.
A controlled re-derivation of the strongest prior generated clip measured 30.34 dB on its held-out split. Later candidates use this value only when the evaluation protocol matches.
The strongest generated candidate in this study measured 27.78 dB. Direct viewer review remains a separate gate because a reconstruction can score well inside its source arc and fail as soon as the camera leaves supported coverage.
Choose the accepted front appearance, the geometry-propagated surface study, the held full-turn research base or the playable great room. Each state opens in the canonical Studio with its camera, review scope and evidence attached.
Camera, Render, Edit and Evidence operate around the same world while the visible representation changes.
Choose a Studio state
Current bounded exterior appearance, reviewed across a six-degree front arc.
Viewing bounds are based on direct review of 52 renders: four models at thirteen azimuths in 10° steps, each marked clear, degraded or gone. Two automated alternatives were rejected. Sharpness rises when Gaussians stretch and foliage fragments. Solved camera positions describe capture locations; direct review establishes the readable range.

Across the four models, readable degrees gained per camera degree fall away steadily: 5.06× readable degrees per camera degree, then 1.75×, then 1.01×, then 0.83×. A 10× spread in capture width bought 20° of additional readable view, and the two widest captures return their own solved arcs and little more. Beyond a point, the win has to come from source appearance and reconstruction quality.
The current bands were generated from the sealed viewer-arc-20260811-r2 manifest. The replaced instrument and its full-resolution frames remain in the archive as rejected range evidence.
Feeding generated frames into a reconstruction costs −1.43 dB against a baseline built from solved frames alone, measured on held-out views. A warped frame mixes resampled source pixels with invented disocclusions, and the warp records that provenance per pixel, which suggested marking invented regions transparent so the trainer could ignore them. A controlled test punched a rectangle to transparent in half the training views to see which way the trainer reads transparency; the other views and all training settings stayed untouched.

So transparency means empty space to this trainer, and masking real disocclusions that way deletes geometry other views actually support. Two short training runs completed in about four minutes. The result ruled out a larger per-pixel masking experiment. Whole-frame selection by invented-pixel fraction remains untested.
The controlled result is preserved as refuted in alpha-semantics-20260811-r1. It stopped the larger per-pixel masking experiment before additional training.