A comparison in today's sweep puts Gendo, ChatGPT, ArchiVinci, and Spacely AI against the same SketchUp input. That is already more useful than four unrelated vendor demos. A shared source view controls the building and camera. It does not control how many generations were attempted, which prompt was rewritten, which failure was discarded, or why the selected frame survived.

That omission matters because an architect does not buy the best isolated image. A studio buys a probability of reaching an acceptable image before the review starts. One polished winner cannot show that probability. A contact sheet can.

One image hides the distribution

Suppose Tool A produces one excellent frame, one usable frame, and six failures. Tool B produces six usable frames and two dull ones. A conventional roundup can crown Tool A by publishing its exceptional frame. A deadline-driven team may rationally choose Tool B because the result is easier to repeat.

This is not a claim about those four named products. Their official descriptions confirm different interfaces and purposes. Gendo describes a collaborative canvas for sketch upload, rendering, material changes, and iteration. OpenAI documents general image creation and editing from uploaded images, including selections that may extend beyond the highlighted area. Those are not identical operating contracts, even when both can return a convincing architectural picture.

The contact sheet makes variance visible without inventing a single score. Show every output generated under the declared budget, in order, at equal size. The viewer can see whether facade rhythm holds, whether glazing changes between attempts, and whether the selected image is typical or lucky.

The rejection pile is not behind-the-scenes material. It is benchmark evidence.

Fix a budget before the first generation

Do not give every tool one attempt. Some products are designed around quick branching, others around controlled editing. One-shot testing can reward a good default and punish an interface built for refinement. Instead, fix an equal evaluation budget that resembles practice.

A useful small test allows eight initial generations and two targeted revisions per tool, with a 30-minute operator limit. If pricing or rate limits make equal generation counts unrealistic, fix the spend and the time instead. State which resource was held equal. Do not quietly grant more reruns to the reviewer’s familiar tool.

Freeze the source view, image dimensions, required brief, forbidden changes, and acceptance checks before starting. Record tool version or access date because hosted services change. Preserve the original prompt and every rewritten prompt. If a platform automatically modifies inputs or offers presets, name the preset.

PublishMinimum recordWhat it reveals
Source panelUncropped input, resolution, cameraWhat geometry the tool actually received
Full contact sheetEvery output, chronological orderVariance, repeated defects, lucky frames
Prompt trailInitial prompt plus each revisionOperator steering and tool-specific coaching
Selection marksKeep, repair, reject, with reasonThe rule used to choose the winner
Budget lineTime, attempts, credits or costChance of success under a real constraint
Final redlineRequired geometry changes markedBeauty separated from project fidelity

Write the rejection rule first

A comparison becomes elastic when rejection happens by taste after the images arrive. Set hard architectural failures in advance. Reject a frame if it changes the number of bays, moves the roof edge, breaks a required railing, shifts the ground line, invents an entrance, or alters an area outside a requested edit. These are examples, not universal rules. The project decides what cannot move.

Then add a repair category. A frame may preserve the building yet need a small sky cleanup or planting mask. Mark it repairable and estimate the correction minutes. Do not mix repairable frames with structural failures. A strange cloud and a changed stair are different kinds of problem.

Finally, identify the keeper rule. Is the winner the first acceptable image, the best image within the budget, or the image needing the least repair? Each answers a different procurement question. The first tests speed to utility. The second tests maximum attainable quality. The third tests downstream labor.

Make operator influence visible

Current tools do not receive the same quantity of evidence. A SketchUp-connected renderer can use model context that a general image editor does not have. A canvas product can accumulate references and edits. A chat interface can accept a source image and plain-language corrections. Fairness does not mean pretending these differences vanish. It means showing them.

Capture a simple action log: upload, preset, prompt, generation, selection, mask, revision, export. Note where the operator used prior product knowledge. A novice result and an expert result are both legitimate, but they should not carry the same label. If one tool receives a vendor tutorial and another is used cold, the comparison measures onboarding as much as rendering.

Yesterday we argued for an intervention ledger that counts work between input and winner. The contact sheet is its visual counterpart. The ledger shows labor. The sheet shows the outcome range. Together they prevent a selected frame from erasing the path that produced it.

Read patterns, not rankings

Once every frame is visible, the useful questions change. Does a tool repeatedly reinterpret glass as void? Does it preserve the camera but vary materials? Are trees consistently pasted across railings? Does the first result usually work, or does quality jump only after a carefully masked edit? Patterns give a studio something it can plan around.

Report acceptance rate as a count, not a spurious precision score. “Five of eight initial frames kept the opening count” is clear. “Geometry fidelity: 8.7” conceals the measurement. Add median time to the first acceptable frame and total repair minutes for the selected output. If the sample is small, say so.

A public review can remain compact. Put the source first, the complete thumbnail grid second, and enlarged keepers below it. Link prompts and action logs in a plain table. The page does not need laboratory theater. It needs enough evidence for another architect to understand what was selected out.

Keep the thumbnails unedited. Do not color-correct one tool, crop away a broken corner, or enlarge only the favorite before the first comparison view. If a platform returns a different aspect ratio, place the full output inside a common frame and label its native dimensions. Retouching can appear later as a documented workflow stage. The first sheet should preserve what actually arrived.

Archive the batch with stable filenames that connect each image to its prompt and action row. A folder named by project, tool, access date, and attempt number is enough. This small discipline lets a team revisit the evidence after a hosted product update instead of relying on screenshots detached from their settings.

Our take: the winner is not enough

Same-input comparisons are moving in the right direction. They answer a basic problem with vendor galleries: unrelated projects make direct judgment impossible. The next improvement is to stop treating output selection as neutral.

A renderer that produces a brilliant frame once may be right for concept exploration. A steadier tool may be right for tomorrow's client set. Both conclusions can survive on the same contact sheet. Neither can be defended by hiding seven frames.

Run the batch. Mark the failures. Put every square on the page.


Editorial basis: the 13 September 2026 ArchiGen AI intel sweep, including a same-SketchUp comparison of Gendo, ChatGPT, ArchiVinci, and Spacely AI. Product interaction claims were checked against current official Gendo and OpenAI pages. This article proposes an evaluation method and does not claim hands-on testing or rank the named tools. ArchiGen AI carries no sponsored placements.