What input should architects use to benchmark an AI renderer?

Use the weakest source image your normal workflow still considers acceptable, then hold the camera and key design facts fixed. Test one first pass and one revision. A polished demo input measures peak output; an ordinary flawed handoff reveals geometry drift, cleanup time and the true cost of an accepted image.

A gray viewport export lands in the folder at 4:47 p.m. The glazing is dull. Two entourage figures clip a planter. The facade materials are technically different and visually identical. Tomorrow's presentation needs atmosphere, not a redesign.

This is the input an AI renderer should meet during a trial. Vendor galleries prefer the disciplined clay view with clean edges, settled materials and a camera chosen by someone who knew the final crop. Of course they do. A violin also sounds excellent after tuning.

Today's market sweep surfaced an unusually good instruction in Rendero's comparison of 11 AI rendering products: test with the weakest input your team normally receives. The page does not report that test across the products. It gives us the right place to begin one.

Why does a polished input hide the bill?

Every missing decision gets made somewhere. If the source does not distinguish spandrel from vision glass, the model may invent a distinction. If planting reads as green foam, it may replace species and location together. If a reflected ceiling plan is absent from the view, the rendered soffit can become architectural fan fiction.

The output may still look excellent. That is the trap. Visual appeal can rise while design fidelity falls, and the repair arrives later when someone compares the image with the model. Our geometry hallucination checklist treats each changed opening or junction as an observed defect, not a vague feeling about realism.

Do not ask whether the image is impressive. Ask who must repair the decisions it made without permission.

A weak-input benchmark makes that labor visible during selection. It also exposes a product's preferred source. One tool may recover a loose sketch gracefully but smear a low-contrast model export. Another may preserve a model view and produce wooden planting. Neither result makes the product universally good or bad. It tells you what the tool needs from the handoff.

What counts as weak but acceptable?

Weak does not mean broken. A source with the wrong camera, missing massing or unapproved openings cannot establish whether the renderer preserved the design. The benchmark begins at the studio's actual acceptance line: enough information to identify errors, with enough roughness to demand useful work.

ConditionKeep in testReject before test
CameraAwkward crop or plain compositionWrong viewpoint for the intended decision
GeometrySimple surfaces and unresolved detailMissing or inaccurate primary form
MaterialsFlat placeholders with clear boundariesOne material ID covering unrelated elements
EntourageBasic planting and people for positionObjects that obscure the area being judged
ResolutionNormal internal review exportCompression too severe to read edges

The base-render acceptance gate draws the useful line. Upstream design errors go back to the source. Downstream presentation defects enter the AI test. Mixing them rewards the renderer for concealing bad information.

Build five frames, not fifty prompts

The benchmark needs one source and four outputs. More generations make a prettier contact sheet while quietly giving the operator more chances to curate luck.

Frame 0: the accepted source

Save the file exactly as received. Record dimensions, application, export settings and date. Mark five to ten invariants on a copy: camera crop, opening count, roof edge, structural rhythm and material boundaries. The marks create a review contract.

Frame 1: the default first pass

Use the simplest documented workflow and a literal brief. Do not spend half an hour translating the building into model folklore. If the product requires specialist prompting to survive an ordinary source, that training burden belongs in the result.

Frame 2: one controlled repair

Request a local correction to one material, planting zone or patch of light. Keep the accepted first frame as the starting point. This tests whether revision is an edit or another roll of the entire image.

Frame 3: delivery output

Export or upscale to the size your presentations require. Inspect thin rails, window divisions, signage and repeated textures at full size. Our native-resolution audit explains why the upscaler belongs inside the comparison rather than after it.

Frame 4: the difference image

Align the delivery output with Frame 0 and produce a difference matte or rapid overlay. Some movement is intentional because lighting and texture changed. The camera, silhouette and marked invariants should remain legible. A product that cannot keep them still needs a narrower role.

How do you score the result?

Use counts and minutes where possible. Count changed invariants. Time the first acceptable output and the controlled repair. Record generations, credits, exports and manual cleanup. Then note the highest project stage at which the result could be shown without misrepresenting resolution.

Keep aesthetic preference separate. Two reviewers can score appeal without seeing the product name, but appeal should not cancel a moved column. The photorealism defect ledger is useful here because it separates geometry, materials, light and finish instead of turning them into one mushy number.

A simple procurement table needs six columns: accepted on first pass, invariant defects, revision attempts, elapsed minutes, credits consumed and manual repair minutes. Add one sentence on the return path. Does an accepted material decision enter the model, or does it remain stranded in a rendered image?

ComfyUI changes the operator, not the standard

Current community questions ask for ComfyUI graphs that take an existing architectural render to 4K while improving light, texture, reflections and planting. ComfyUI can split those jobs across img2img, masks and spatial controls. Its graph also lets an operator tune around a weak source in ways a one-button product does not.

That flexibility is valuable, but graph-building time counts. Official ComfyUI documentation describes the source image as an additional condition in img2img, while its inpainting workflow limits regeneration to a selected area. Both mechanisms can protect an accepted source. Neither recognizes a mullion as a project commitment. The operator still defines the evidence and checks the output.

If a skilled ComfyUI operator spends ninety minutes preparing maps and masks, compare that labor with the subscription product's cleanup, not with its generation timer. Seconds per image is a party trick. Minutes to an accepted revision is a workflow measure.

When does the weak-input test end?

Stop after the controlled repair and delivery export. The point is to observe a routine handoff, not to discover the maximum quality available after heroic attention. If every product fails, improve the studio's source acceptance line and run the same test again. That is a useful result too.

The best trial image is rarely the one a studio should buy around. Choose the product that survives the file people actually send at 4:47, then returns an image nobody has to apologize for at nine.


Evidence note: this is a proposed benchmark, not an ArchiGen hands-on comparative test. Sources checked 6 October 2026: Rendero's instruction to compare products with the weakest normal team input; current community requests for ComfyUI architectural render enhancement; official ComfyUI img2img and inpainting documentation. No product performance result is claimed.