Why I Start My AI Pipeline With an Intentionally Terrible Image

Wait 5 sec.

Generating a beautiful room with AI is easy now. But good luck generating a photorealistic 360° interior that actually follows the exact floor plan.I had two days to build a presale prototype that could help win a project. I did just that, and its goal was simple: render the rooms of a house as photorealistic 360° interiors, make it possible to show the same spaces in different interior styles, and keep every result faithful to the source floor plan. Windows, doors, hallways, walls, and how rooms connect to each other all had to stay exactly where the floor plan put them, whatever style the room ended up in. In the end, every room had to turn into a 2:1 equirectangular panorama that someone could explore in a 360° viewer.A generative image model made a convincing interior right away, as long as it was a normal flat shot. The trouble started when I asked for the same room as a 2:1 equirectangular panorama. The image is supposed to wrap around into itself, and it didn't. A wall that started on the left edge of the picture wouldn't match up with itself on the right edge. The room looked fine from any single angle. The moment you turned all the way around inside it, the seam gave it away.Having given it a handful of failed prompts, I had a final-render Eureka right in the middle. The working architecture was a hybrid of this formula:Use deterministic code for parts with one correct answer, and generative AI for parts where variation is useful.For this prototype, that meant rendering the geometry first, in the same 360° projection the final image would use, then asking the image model to turn that ugly render into a realistic room.The broader lesson was even more useful: when a model keeps ignoring a constraint, a better prompt isn't always the fix. Sometimes the constraint needs to reach the model in a different form entirely.Photorealism is easy, geometry is hardcoreA normal interior-generation prompt is forgiving: get a window a few inches off, and nobody notices. But a floor-plan-driven, 360° experience is different. The panorama has to close into a loop: the wall on the far left of the image and the wall on the far right are the same wall. That's exactly where the model struggled. It could place a fireplace on the right wall without any trouble in a flat image. Ask for the same room as a panorama, and that same wall might not agree with itself at the seam.My first attempt was the obvious one: give the model the plan, describe the desired style, and ask it to produce the room as a 2:1 equirectangular panorama.The result looked excellent at first glance and was absolutely wrong structurally. A top-down plan and an eye-level 360° panorama represent space in two completely different, unalike ways. In one step, the model had to infer 3D geometry, pick a camera, preserve topology, understand openings, project the scene into an equirectangular image, then furnish and style it.Why would metric correctness win that fight?3 approach experiments that clarified the architectureEach subsequent iteration failed in a different way, pointing to a task the model shouldn't have been handling in the first place.Attempt One: Six Faces, Zero AgreementInstead of asking the model to create an equirectangular image directly, I generated six 90° views and converted the cubemap into a panorama with deterministic projection math.That fixed the output projection problem completely: it was now valid by construction.But the six faces did not agree with one another. Furniture duplicated across boundaries, object positions drifted, and the seams exposed the fact that each face was still an independent generative decision. A shared text layout reduced some duplication, but consistency remained a soft instruction.Attempt Two: Seamless, but Clueless About the HouseA model built specifically for equirectangular images handled the wraparound naturally. The panoramas came out seamless, but generic, because the model still had no real sense of this particular house's geometry.This distinction mattered: correct projection does not imply correct geometry.Attempt Three: Reverse-Engineering What I Already KnewI also tried a different route: estimate depth from one generated image, turn that into a point cloud, then reproject it from a new camera angle.This was much more fragile than it looked on paper. The depth estimate was only relative, not exact measurements. I didn't know the original camera settings. Windows looked like holes leading into some distant exterior. Rotating the point cloud exposed gaps and visual artifacts wherever a surface was missing.That experiment flipped the problem around for me. I already had the floor-plan geometry; why then was I asking a model to rediscover it from an image?If a sub-problem has a closed-form answer, the model should not be approximating it.Architecture that worked: render geometric scaffold firstThe real breakthrough came from an unrelated corner of the internet. X and other platforms were full of challenges where people fed a generative model a photo of a person, a room, anything, and the model repainted it while keeping the original scene intact. Only the style changed.That gave me an idea. Instead of asking a generative model to draw a 360° room from nothing, what if I handed it an already-correct, schematic 2:1 panorama and asked it to repaint that? Add materials, lighting, furniture, make it look real, and leave the geometry untouched.The first version worked. Earlier, when I'd asked a model to draw a room directly as a 2:1 panorama, it produced something, but never a whole, coherent image. This time it did. And the point wasn't that this version drew a nicer room, or even that it respected the floor plan. A JSON-driven prompt could probably respect the floor plan too, and make a decent-looking room while it was at it. But the real win was what this approach let me skip: a full, production-ready 3D build of the house, with real models, materials, and lighting, just to get a floor-plan-accurate 360° view.The final prototype started from structured whole-house JSON. Rooms were axis-aligned rectangles with ceiling height and openings on their walls. Each opening had a type (window, door, or open passage) + position, dimensions, and what it connected to.Before rendering, I validated the geometry. Shared walls had to align, and interior openings had to exist consistently on both sides of the connection. That let me treat the house as one scene rather than a set of unrelated rooms.Each room then became an intentionally minimal 3D shell: floor, ceiling, and planar wall quads split around openings. The scene only needed to encode structure, so I left secondary elements (furniture, textures, lighting, etc) for later generative steps.A small software ray caster rendered an eye-level 2:1 equirectangular panorama. For every output pixel, the renderer cast a ray into the scene, found the nearest intersected surface, and assigned a simple semantic color. Rays could pass through open portals into neighboring rooms and out through exterior windows.whole-house JSON      ↓geometry validation      ↓minimal 3D scene for the whole building      ↓equirectangular ray cast      ↓geometric scaffold / conditioning image      ↓generative image restyle      ↓photorealistic 360° panoramaAt 1024×512, the shell rendered in roughly 0.3 seconds on CPU; at 2048×1024, it took about 3 seconds. The geometry pass was cheap compared with the image-model call.The render was intentionally ugly: its job was to make the structure difficult to misunderstand.Portals were more important than realismRendering the whole building was a major improvement over rendering each room as a closed box.If the kitchen opened into the living room, rays passed through the opening into the actual living-room geometry. If there was a window on the far wall of that neighboring room, it could be visible through the portal.Adjacency was no longer something the prompt had to describe. It was already visible in the image.That changed the image model's task from:understand this floor plan, reconstruct the space, choose the right projection, and make it beautifulinto something much narrower:keep this structure and replace the crude surfaces with a realistic interior in this styleFor a prototype, that was night and day, especially next to where this whole experiment started.3 Lessons From Trading JSON for a PictureThe scaffold's geometry wasn't any more accurate than the plan's. What changed was the form the model had to read it in.With JSON, the model has to build the whole 3D scene in its head from raw numbers. A floor plan is a bit closer, but the model still has to picture it rotated from a top-down view into an eye-level one. The scaffold skips both steps. It already looks like the final output: same angle, same projection.All that trial and error finally bore fruit. Three practical lessons so far:More context was not always betterAt one point I supplied both the rendered scaffold and the original plan. Adherence became worse: the model tried to reconcile 2 representations and sometimes invented structure in the process.Removing the plan and keeping only the rendered shell improved the result. Redundant constraints don't automatically reinforce each other, especially when they arrive through different modalities.Style language can accidentally become geometryA style prompt that mentioned things like a stone fireplace or archways caused the model to insert those objects even when the shell showed a solid wall.But if the geometry image already shows where things are, the text shouldn't compete with it, even indirectly.  So I fixed it by keeping the style prompt mostly about materials, palette, and lighting rather than architectural objects.Constraint encoding mattersClosed doors were a good example. Rendering the door area as a plain wall sometimes caused the model to invent a door elsewhere. Rendering it as an empty opening suggested a corridor. Rendering it as a distinct opaque door panel in the correct location produced much better behavior.The source fact was the same in every case. The visual encoding changed how the model interpreted it.This is why I would not describe the system as prompt engineering. The most important engineering happened before the prompt, in the interface between deterministic geometry and probabilistic generation.Bugs Got Smarter, Not FewerThe scaffold made the system much more reliable, but it did not turn an image model into a CAD renderer. The conditioning was still soft.In one bathroom test, the door stayed roughly where the geometry placed it, but a small high window became much bigger, and the model added extra structure that wasn't there. Other times, it swapped one type of opening for another, or added a door in a solid wall.What's interesting is that the mistakes got smarter. The model almost always kept the overall wall layout and the curve of the panorama right. What it got wrong was more semantic: was this a window or a glass door, a passage or a doorway, a wall or a niche.The prototype also cut some corners on purpose. Each room was generated on its own, so furniture seen through a doorway might not match how the next room actually looked. Views outside the windows were made up. And there was no automatic checker for openings, no way to fix one small part by hand, and no proper retry or save system.Those are fixable, more familiar problems: shared context between rooms, extra signals like depth maps, a way to redo one bad spot instead of a whole room.But really, I was going for something much simpler here: prove that a convincing, floor-plan-accurate 360° experience was possible without first building an expensive, fully authored 3D modeling and rendering pipeline for every room. The prototype came together in a few hours of hands-on iteration spread across 2 days, enough to move the idea into MVP development.Hybrid is a bridge, not a destinationWithout abandoning this approach, the first thing I'd strengthen is the interface between the deterministic and generative parts.The renderer already knows the real geometry, so besides a color scaffold, it could also output depth, normals, semantic masks, and opening masks. These extra signals would give the model stronger structural guidance. The same geometry could also power a verifier that checks whether generated openings match the real scene and retries the ones that don't.But once the product proves demand, there's a second path: move even appearance back toward deterministic 3D.For an MVP, letting AI handle appearance makes sense: drawing the geometry is cheap. The real cost in a full visualization stack is everything else: the asset library, furniture placement, materials, lighting, scene authoring, rendering infrastructure, editing tools, and people's time.Once the product earns that budget, those same costs buy real guarantees: real assets, controlled lighting, and hand edits instead of leaving it to a model. Generative AI can still help you with ideation, texture exploration, or style suggestions. It just no longer has to own the final appearance.That is an important product-engineering distinction: the hybrid was the cheapest architecture that could answer the next product question, a stepping stone rather than the final form.Broader lesson: move deterministic boundary deliberatelyThis pattern applies far beyond interior visualization.A lot of AI systems get unreliable because we ask one model to solve several different kinds of problems at once, including parts that regular code could already solve exactly. Teams then try to patch that with longer and longer prompts.A better approach is to ask what kind of uncertainty each piece actually has.Give deterministic code the parts with one right answer: geometry, topology, coordinate math, projections, validation, calculations, permissions, anything you can test.Give generative models the parts with no single right answer: appearance, language, style, texture, ideas, anything with room for many valid results.And pay just as much attention to how the two connect.In this prototype, each useful iteration moved another responsibility from probability into code: projection, wraparound, geometry, adjacency, and previews. Appearance stayed generative not because it has to be, but because building it by hand (real assets, furniture placement, materials, lighting, a full 3D pipeline) was the expensive part, and we chose to put that off until the product had earned the time and budget for it.The principle I would carry into other AI products is simple:Render the constraints, generate the rest - and keep moving that boundary as the product earns the budget for stronger guarantees.