Using AI /
How to prompt AI image models with reference images
Most people get generic AI images because they ask the model to invent too much. They type "create an epic product photo of a shoe", then wonder why the result looks fake, generic, like the same slop everyone else is generating. The fix is almost dull. Stop asking the model to imagine your image. Show it one that already looks the way you want, and build the prompt around that.
The model is not a creative director
A model writing an image has no taste. It does not know your brand, your lighting, the camera language of a good campaign, or what "premium" means in your category. So it guesses, and its guess is the average of everything it has seen. The average is the slop.
A reference image fixes this in one move. It hands the model composition, lighting, framing, mood, texture, camera angle and the general sense of what good looks like, without you describing any of it in words. A screenshot is worth more than a paragraph of adjectives.
Find the reference before you write a word
Start by hunting for an image that already feels close to what you want. Pinterest, brand sites, product pages, campaign shots, editorial layouts, old ads, Midjourney galleries, landing pages.
Do not only look in your own category. The visual language travels.
- A running shoe ad can teach a bike ad.
- A car launch can teach a product hero shot.
- A watch campaign can teach the lighting for a supplement bottle.
- A hiking homepage can teach a trail running image.
You are not copying the picture. You are showing the model the grammar.
Give each reference a job
Image models blend whatever you give them. Upload three pictures and say nothing, and you get a smoothie. So label them.
- Image 1 is the style. Mood, lighting, composition, camera, framing, typography feel.
- Image 2 is the product. The exact thing that has to stay accurate.
- Image 3 is the detail. A logo, a material, a texture, a sole, a piece of packaging.
Now the model knows what each picture controls, instead of averaging them into one.
Hard lock anything that has to stay real
This is the move that separates usable product images from the uncanny ones. Without it, the model quietly redraws your product. Bikes get the wrong geometry. Shoes grow strange soles. Packaging text turns to nonsense. Logos go fake.
So say it plainly in the prompt.
- "Hard lock the product from Image 2. Keep it 100 percent identical."
- "Do not change the geometry, proportions, colours, branding, logo placement, materials, reflections, hardware or text."
- "Only change the background, lighting and treatment."
The model respects a constraint it is told to respect. It invents anything you leave open.
A few habits that help
None of these are clever. All of them save you a bad afternoon.
- Build the prompt in one chat, then generate in a fresh one. The building chat fills up with old corrections and rejected ideas that bleed into the image. Start the generation clean. New chat, upload the references, paste the final prompt.
- Use the strongest model and the highest quality setting. Instant modes are fine for rough ideas, not for the final thing. The model needs room to hold the references, the constraints and the accuracy all at once.
- Do not regenerate forever. One focused retry is fine. After that, long chains start to rot, with repeated artefacts, warped text and an over-cooked look. When it goes wrong, do not push the same chat harder. Work out what failed, tighten the prompt, and start fresh.
The structure that holds it together
Once you have the references, a good prompt is just a checklist. Say what you want. Hand out the reference roles. Lock the subject. Set the format. Give the creative direction. Place things in the frame. Describe the light and the lens. Spell out any text exactly. Set the physical rules. End with what to avoid.
That last part earns its place. An avoid list is where you head off the obvious AI tells, the floating parts, warped logos, plastic skin and fake sale badges, before they happen.
Here is the skeleton I reuse.
What it looks like filled in
Here is the same skeleton turned into a real product shot. Notice how much of the work is constraint, not description.
The same skeleton flexes to a 9:16 ad, a vertical infographic, an action shot or a website hero. You change the format line, the creative direction and the text. The reference roles and the hard lock stay exactly where they are.
The rule
If the image matters, do not start with the prompt. Start with the reference.
The reference teaches the model what good looks like. The prompt tells it what to keep, what to change and what to avoid. Almost everything generic about AI images comes from skipping the first step and hoping the second one carries it on its own.
