I Kept Getting Bad AI Art From Great Reference Photos—So I Built My Own Image-to-Prompt Generator
A few months back I had this photo saved on my phone. Soft window light and muted blue-grey tones, that slightly moody, cinematic look you see in a lot of good product shots. I wanted to recreate something with a similar feel for a client project using Midjourney.
So I typed what I thought was a decent prompt. "Moody photo, blue tones, soft light." The result looked nothing like the reference. I tried again. Added more adjectives. Still off. After about the sixth attempt I finally admitted the problem wasn't Midjourney—it was me. I could see what made the photo work, but I couldn't put it into the right words fast enough.
That's when I went looking for an image-to-prompt generator to do the heavy lifting for me.
The Problem With Most Image to Prompt Tools:
I tried a handful of them. Some were fine, but almost all of them had the same issue: you'd upload a photo and get back one flat sentence. Something like "a woman standing near a window. "Technically true. Completely useless if you actually want to rebuild the lighting, the colour mood, or the composition.
A couple of others worried me for a different reason—you had no idea where your image was actually going once you hit upload. For client work, that's not something I wanted to gamble on.
So instead of hunting for the "right" tool, I just built one for myself. What started as a weekend project turned into something I now use almost daily, so I figured I'd share how it actually works—partly because I think it'll save someone else the same headache I went through, and partly because the mistakes I made building it taught me more about prompting than any tutorial did.
What This Image to Prompt Generator Actually Does:
Instead of guessing at a description, the tool breaks your image down into the stuff that actually matters for an AI prompt:
All of that gets translated into plain English descriptors and assembled into a prompt. And here's the part I'm genuinely proud of: none of this happens on a server. The whole analysis runs right there in your browser tab. Your image never gets uploaded anywhere, which was non-negotiable for me once I started using it for client references.
How to Actually Use It (Step by Step)
Mistakes I Made Building This:
My first version only had one output style, tuned for natural language. I quickly realised that if you feed a Midjourney user a paragraph of prose, they have to rewrite half of it anyway. That's what pushed me to add separate formatting for each target model instead of pretending one prompt style fits everywhere.
I also originally cranked every prompt up to maximum detail by default. Turns out that's the wrong move for something like Midjourney, where a long, over-specified prompt can actually fight against the model's own style. That's why the detail levels exist now — Simple, Balanced, and Detailed aren't just a nice-to-have, they came from actually seeing bad output and working backwards to figure out why.
Who This Is Actually Useful For:
If you've ever typed "aesthetic photo, nice lighting" into an AI generator and gotten something completely unrelated back, you already know the gap this fills. It's not really an image describer in the boring, one-line sense — it's closer to reverse-engineering the actual visual ingredients of a photo so you can rebuild that feeling somewhere else.
A Few Tips From Using It Myself:
1. Use a source image with a clear subject and decent lighting—a dark, blurry photo gives the analysis less to work with, and the prompt will reflect that.
2. For Stable Diffusion, always use Detailed mode. The negative prompt it generates alongside the main one genuinely cuts down on the weird artifacts and bad-anatomy issues.
3. If your first result feels close but not quite it, don't start over—hit Variation two or three times before you touch the settings. Half the time the phrasing was the actual issue, not the analysis.
4. Match your aspect ratio. The tool reads your image's actual proportions and builds that into the Midjourney parameters automatically, so your output framing tends to match your reference a lot more closely than a random square generation would.
