The Technical Story Behind InstaRoom’s Evolution: From Four AI Models to Two API Calls

In my last article, I wrote about what ikebana taught me about building InstaRoom, how stripping away features to focus on one use case made the product better. But I didn’t get into what changed under the hood. This is that story.

Live: instaroom.ai

When I first built InstaRoom, Stable Diffusion was the state of the art for image generation. And it was impressive, but also completely uncontrolled. Upload a living room photo, ask for a boho redesign, and the AI might move your windows, delete a wall, or float a couch in midair. Beautiful output. Wrong room.

Four models working together

The fix was ControlNet, not one model, but four working together. Canny for edge detection to preserve architectural lines. Depth for spatial awareness so furniture stayed on the floor. MLSD for straight-line geometry. Normal maps for surface orientation and lighting.

I spent weeks tuning the balance between these. Too much control and nothing transformed. Too little and the room structure fell apart. We landed around 0.7 strength and 7-8 guidance scale, the sweet spot where the style changed but the bones of the room stayed intact.

The infrastructure matched the complexity. Two UI servers behind a load balancer, a separate API server coordinating with Replicate for AI inference, Postgres for tracking, S3 for storage. When a user clicked “Redesign,” the request bounced between five services: the API server would fire off to Replicate, Replicate would send progress callbacks as it worked (10% done, 20%…), and the UI server polled the database every ten seconds waiting for the result.

It was a lot of moving pieces. And it worked: 250,000+ generations at 99.7% uptime. I’m proud of what V1 became, especially the interface. For a design product, how it looks and feels matters as much as what it does, and V1 felt like a real professional tool.

What the usage data kept telling me

Most people just wanted a simple room refresh. And they kept asking the same thing: “this looks great, but where do I buy this stuff?” With Stable Diffusion generating pixels, not product knowledge, I didn’t have a good answer.

Then Gemini launched with native image generation, one model that could reason about text AND produce images. That wasn’t just a better model. It was a different capability entirely, and it made a different product possible. Instead of generating a beautiful room and leaving users to figure out what’s in it, I could now start with the shopping list.

How V2 works

Step 1: You upload a room photo and say something like “refresh this in boho style.” Gemini analyzes the room and generates a specific shopping list: a jute area rug from Wayfair, a fiddle leaf fig from Pottery Barn, abstract wall art from Target, linen curtains from IKEA. Specific items, specific prices, specific retailers. All under $1,000.

Step 2: That shopping list feeds directly into a second Gemini call for image generation. And here’s what makes it different from V1: the visualization actually shows those specific items in your room. The throw pillows in the rendering are the throw pillows on the list. The lamp you see is the lamp you can go buy. Shopping list and visualization connected, not two independent outputs.

The architecture went from five distributed services to Docker Compose on a single server: React, FastAPI, PostgreSQL, two Gemini API calls in a background thread pool. Simpler to build, simpler to debug, simpler to maintain.

I’m still proud of V1‘s polish. The conversational chat interface in V2 doesn’t have that same visual quality yet. The best AI products need both: the right solution AND the experience to match.

The bigger lesson

The AI landscape moves fast enough that the right architecture today might not be the right architecture in six months. V1’s multi-model ControlNet pipeline with distributed async callbacks was the right call when I built it. V2’s two-step Gemini pipeline is the right call now, not because V1 was wrong, but because new capabilities made a fundamentally different product possible.

Building AI products means paying attention to when the ground shifts, and being willing to rebuild around a better problem, not just a better model.

About the author

madhavi

Add Comment

By madhavi

madhavi

Get in touch

Solutions architect and forward deployed engineer in the San Francisco Bay Area, building production AI systems. Previously Oracle, Apple and Yahoo.