Put a Veras viewport beside a Midjourney image and a ComfyUI enhancement, then ask which looks best. It feels like a comparison because three pictures sit in a row. It is actually three different assignments. Veras may have been told to respect a Revit model. Midjourney may have been given a sentence. ComfyUI may have received a nearly complete V-Ray frame plus depth and edge maps. The result with the most attractive light may also have started closest to the finish.
Architecture teams should care because input determines risk. A tool that consumes a live BIM view inherits measurable geometry. A tool that consumes a raster view sees pixels, not walls. A tool that starts from text receives no project truth unless the prompt can somehow describe every bay, sill, setback, and material joint. Ranking the outputs without accounting for that difference rewards invention and calls it performance.
The input is the tool’s contract
Every renderer makes a quiet contract with its user: give me a certain kind of evidence and I will preserve some of it. The useful question is not whether the output is photoreal. It is which evidence the tool accepts, what it promises to retain, and what the operator must rebuild when it fails.
A BIM-integrated renderer accepts camera, geometry, materials, and often object identity directly from Revit, SketchUp, Rhino, or Archicad. Its contract is strong on form and convenient iteration. A general image generator accepts text and reference images. Its contract is strong on mood and weak on exact dimensions. A controlled diffusion graph can accept a rendered view, mask, depth map, edge map, normal pass, prompt, seed, and model. Its contract can be very specific, but the operator has to assemble and maintain it.
An enhancer accepts an image that already contains most design decisions. It may be excellent at surface texture, planting, light, or apparent resolution. That does not make it a renderer from BIM, and it should not lose points for failing a job it never claimed. Categories matter because evidence matters.
A beauty shot tells you what a tool can invent. An input test tells you what it can be trusted to keep.
Use one project and four evidence levels
A fair studio evaluation starts with one small project that the team knows well. Choose a view with repeated windows, a visible corner, one thin element such as a railing, transparent glazing, and a material transition. These features expose drift quickly. Freeze the camera and define four evidence levels.
Level one: text only
Give each compatible tool a concise brief: building type, massing, primary materials, weather, time of day, lens, and viewpoint. This tests concept generation. Do not score dimensional fidelity because none was supplied. Score composition, relevance, prompt responsiveness, and the number of usable directions produced in fifteen minutes.
Level two: one reference image
Supply the same clay viewport at the same resolution. Disable extra control inputs. This tests how well each product reads visible form from pixels. Score silhouette, opening count, camera stability, and whether materials land in the requested regions. Record the strength setting because a tool can preserve form by making almost no change.
Level three: geometry evidence
Now give compatible tools the strongest structure they accept: a live model, exported depth, edge map, normal pass, or segmentation. This is the working-architect test. Score roofline, facade rhythm, floor count, mullions, railing continuity, and boundary placement. A tool that cannot accept structural evidence is not penalized universally. It is marked unsuitable for geometry-critical work.
Level four: local correction
Ask for one precise change: replace only the ground-floor cladding, change only the sky, or add planting behind the railing without moving it. Use a mask when supported. Score collateral change outside the edit zone, time to a keeper, and whether the next view can receive the same correction. This level often reverses the ranking from the first beauty shot.
| Evidence | What it tests | Do not pretend it tests |
|---|---|---|
| Text brief | Ideas, mood, composition | Design fidelity |
| Clay viewport | Image guidance, camera retention | Object-level understanding |
| Model or control pass | Geometry retention | Final polish by itself |
| Mask and correction | Local control, iteration cost | First-pass speed |
Score survival, not just shine
The final sheet should include two scores. The first is visual quality: light, material credibility, detail, and composition. The second is evidence survival: how much supplied project truth remains correct. Keep them separate. Combining them into one number hides the trade that architects most need to see.
Count specific defects instead of assigning a vague fidelity score. Did the tool change the number of window bays? Did the parapet move? Are handrails continuous? Did glazing become an opening? Did brick cross a material boundary? Can the same seed or settings carry to a second camera? A five-minute redline gives a more honest result than a panel debating whether one sunset feels cinematic.
Track operator time too. Include setup, generation, selection, correction, export, and the time required to reproduce the result on another view. A fast first image followed by forty minutes of repair is not a fast workflow. Neither is a perfect node graph that takes a specialist half a day to install on a colleague’s machine.
What today’s comparison tables miss
The current crop of articles often compares starting price, BIM integration, speed, and output quality. Those columns are useful, but they treat integration as a checkbox and quality as a property of the final JPEG. Integration changes the quantity and type of evidence available to the model. That changes the assignment itself.
A Revit or SketchUp plugin should be asked how reliably it follows a revised model. A hosted sketch renderer should be asked how quickly it turns a drawing into plausible options. ComfyUI should be asked whether a graph can preserve structure, isolate corrections, and repeat the result. An enhancer should be judged on how much finish it adds without damaging the supplied image. There may be one winner inside each contract. There is no honest overall winner across all four.
Our take
Studios do not need a larger table. They need a test file and a scoring sheet. Use one real project, freeze the view, define the evidence level, and count what survives. Publish the input beside every output when presenting the result internally. If a tool needs a depth pass, show it. If it began with a finished render, say so. If the operator repaired the image in Photoshop, count the minutes.
The best-looking image can still win the concept round. It simply cannot claim it also won the BIM round without taking the same exam. Same building. Same evidence. Then score it.
Written from the 26 August 2026 intel sweep, which surfaced several current AI rendering comparisons alongside community questions about controlled architecture workflows. ArchiGen AI carries no sponsored placements.