AI Image Generation Prompts: How to Write Better Prompts
Learn how to write AI image generation prompts that communicate subject, composition, lighting, atmosphere and intent without drowning the model in adjectives.
AI Image Generation Prompts: How to Write Better Prompts
There’s a particular kind of frustration that comes with AI image generation.
You know exactly what you want. You can almost see it: the shape, the atmosphere, the light, the strange little detail in the corner that somehow makes the whole thing work.
So you type a prompt and press generate. Something comes back that is technically impressive, occasionally beautiful, and completely wrong.
It has the right subject. The right colours. It may even have the right style.
But it isn’t your image.
That’s where prompting gets interesting.
The goal isn’t to find a secret combination of words that forces an AI model to obey you. There isn’t a universal magic formula. The official prompting guides from OpenAI, Google and Adobe differ in their details (some even hand you templates), but they lean the same way: describe the visual result clearly, give the model the context that matters, and refine what comes back on purpose.
In other words, you’re not casting a spell. You’re directing a scene.
Once you start thinking that way, writing AI image prompts gets considerably easier.
Start with the image, not the prompt
Before writing anything, forget about prompting for a moment.
Ask yourself: What am I actually trying to see?
Not “what keywords should I use?” Not “what’s the best prompt format?” And definitely not “how many times should I write cinematic?”
Describe the image to yourself first. Then imagine the AI isn’t an AI at all. You’ve hired a photographer, a production designer, an illustrator or a 3D artist, and you’ve walked into their studio and said:
“I need an enormous abandoned machine inside a subterranean laboratory. It should feel like something humans built for a purpose they no longer understand. The machine is still running. There is a faint violet light coming from somewhere inside it.”
That’s already a useful creative brief. No prompting jargon, no three hundred words. But it communicates an idea, and that’s where you start.
Before you tell the machine what to make, decide what you are actually trying to summon.
Why vague prompts produce strangely familiar images
Consider this:
A futuristic city.
There’s nothing technically wrong with it. The problem is that you’ve left almost every important visual decision open. The model has to decide:
-
what “futuristic” means
-
what kind of architecture exists
-
whether the city is clean or decayed
-
whether we’re on the ground or in the air
-
what time of day it is
-
whether anyone is there
-
what the weather is doing
-
what dominates the frame, and what the mood is
The model will fill those gaps. And when it has that many decisions to make on its own, the result can drift toward familiar visual patterns.
Try giving it a specific situation instead:
A vast vertical city built inside a geological canyon, viewed from a narrow pedestrian bridge hundreds of metres above the streets. Dense layers of old concrete buildings are connected by suspended transit lines and illuminated walkways. Rain is falling through the canyon and collecting on the glass railings. The upper buildings disappear into cloud and industrial haze. The camera is positioned close to the railing, looking diagonally down into the depth of the city.
Now the model has something to work with. Notice what happened: we didn’t make the prompt more impressive. We made the idea less ambiguous. That difference is the whole game.
A prompt is really a list of visual decisions
Every image contains hundreds of decisions. Some matter. Most don’t. Your job isn’t to describe everything. It’s to describe the decisions that matter, and to leave the rest to the model on purpose.
Imagine the image as a room. What is the room? What’s important inside it? Where are you standing? What’s lit, and what’s hidden? What makes this room different from every other room? Answer those, and you’ve done most of the prompting.
Some people call this prompt engineering for images. Fine. It’s mostly deciding what matters.
Give the model something it can actually see
Words like mysterious, beautiful, premium, emotional, luxurious, futuristic and magical carry your intention, but they’re abstract. The model has to translate them into visible things, and it may translate them differently than you would. OpenAI’s own guide makes a similar point about rainy, neon or low-light scenes: specify scale, atmosphere and colour instead of leaning on mood words alone.
Instead of:
a mysterious forest
try:
a dense forest just before nightfall, narrow trunks disappearing into low ground fog, almost no visible sky, a single warm light glowing far between the trees.
Now “mysterious” doesn’t have to do all the work. The mystery is visible.
Instead of:
a luxurious room
try:
a quiet hotel suite with dark walnut walls, pale stone flooring, oversized linen curtains and a single sculptural chair beside a floor-to-ceiling window overlooking the ocean.
Same move. You’ve turned an adjective into things the viewer can see.
The general rule: when possible, translate an emotion into visible evidence. Don’t tell the model something is eerie. Show it what eerie looks like.
The subject needs a situation
A subject floating in isolation feels generic. Compare:
An astronaut on Mars.
with:
A lone astronaut standing beside a damaged rover on the edge of a Martian canyon, one hand resting on the vehicle while a distant dust storm approaches across the horizon.
The second version gives the astronaut a place, a relationship with an object, an action, a sense of scale, and something happening in the distance. Now the image has a moment. An image isn’t a collection of objects. It’s a frozen event.
Think about where the camera is
This is where prompting starts to feel like directing photography. Ask where the viewer is standing. Beside the subject? Looking up? Looking down? Very close, far away, inside the environment, above everything, hidden behind something?
You don’t need technical camera specs to say it:
View the scene from ground level, looking upward.
A distant aerial view showing the entire structure surrounded by the landscape.
The camera is positioned behind the character, looking toward the illuminated doorway.
OpenAI’s prompting guide tells you to set the composition and any important placement up front, and to describe the framing. Even lens terms like “50mm” or “shallow depth of field” are best treated as cues for how the image should look, not as an exact physical simulation. What these instructions really set up are relationships, and relationships often matter more than decorative adjectives.
camera and framing in the Image Prompt Vault
Scale is a secret weapon
One of the easiest ways to make an image feel different is to establish scale. A person standing beside a building tells us something. A person standing beside a building that is three hundred metres tall tells us something else.
A tiny human figure standing at the base of an enormous circular machine, with the machine extending far beyond the top of the frame.
Now the model has a hierarchy. The person isn’t really the subject anymore. The relationship between human and machine is.
Scale helps with architecture, science fiction, landscapes, machinery, monuments, creatures and surreal scenes. If something should feel enormous, don’t just write “gigantic.” Show the viewer what gigantic means.
Light is not decoration
A lot of people treat lighting as an afterthought. It shouldn’t be. Light tells the viewer where to look. It sets the time of day, the atmosphere, the material. It can make the same room feel safe, threatening, sterile, nostalgic or completely alien.
Compare:
A laboratory with blue lighting.
with:
A mostly dark underground laboratory illuminated by narrow strips of cold blue light along the floor, with a faint warm glow leaking from behind the central machine.
The second version creates a visual hierarchy. The floor leads the eye. The machine becomes the centre of attention. The warm light becomes something the viewer wants to investigate. That’s far more useful than “dramatic lighting.”
OpenAI, Google and Adobe all tell you to describe the lighting rather than leave it to chance. The more precisely you say where the light comes from and what colour it is, the more you’re directing.
lighting and visual direction in the Image Prompt Vault
Colour should have a job
Instead of “colourful scene,” think:
mostly charcoal, deep violet and muted silver, with small areas of warm amber light.
Now colour has a hierarchy. It can separate subject from background, foreground from distance, artificial light from natural light, human elements from machine elements, reality from fantasy.
You can limit the palette on purpose, too:
Almost entirely monochromatic grey architecture with a single saturated red object.
That can be far stronger than asking for “vibrant colours.”
Style should clarify the image, not decorate the prompt
You can ask for photography, illustration, painting, 3D rendering, editorial imagery, architectural visualisation, technical illustration, surrealism, abstract art. But be careful about stacking them. Plenty of AI art prompts online look like this:
cinematic cyberpunk surrealist editorial photography 3D render anime concept art
What exactly are you asking for? That may be a collision rather than a direction.
Decide which visual language matters most, then say it plainly:
Photorealistic architectural visualisation with physically believable materials and subtle atmospheric haze.
Hand-painted surrealist illustration with visible brush texture and muted colours.
Adobe’s Firefly guidance leans the same way: be specific about subject, action, setting, lighting and mood, and use reference images and Firefly’s own controls when words aren’t enough.
Stop worshipping the word “cinematic”
“Cinematic” isn’t a bad word. It’s a lazy one. It gets stuck onto so many prompts that on its own it tells the model very little about what you actually want. A prompt can say “cinematic” and still have terrible composition, boring lighting and no visual hierarchy.
If you want the feeling, describe what creates it:
Wide framing, strong foreground depth, low camera position, atmospheric haze separating the distant structures, controlled highlights and a limited colour palette.
Now you’re describing the ingredients. The word can stay. It just won’t be carrying the whole image on its back.
Negative instructions: useful, up to a point
Sometimes you do need to say what shouldn’t be there:
No text. No logos. No people. No floating objects.
If a product shot can’t have branding, say so. If you’re editing an existing image, “change only the background, and keep the subject, pose, clothing, lighting and camera angle” is a negative instruction wearing a polite coat.
Just don’t build a fifty-item block of everything you can think of. If the model has no reason to go near something, you don’t need to mention it.
Tools also disagree here. Some have a dedicated negative-prompt field. OpenAI’s guide is happy with plain exclusions like “no watermarks” or “no extra text.” Google’s Gemini guidance suggests a different habit where possible: describe what you want instead of what you don’t, so rather than “no cars,” describe an empty, deserted street. Check how your own tool behaves before importing a workflow from another one.
Protect the important idea. Don’t micromanage the universe.
When you already have a good image, stop regenerating
This is one of the biggest differences between beginners and experienced users.
A beginner gets an image that’s 70% right and thinks, I’ll run the whole prompt again. That can wreck the good 70%.
Instead, correct one thing:
Keep the composition and subject unchanged. Make the lighting warmer and reduce the background clutter.
OpenAI’s guide says much the same: ask for one change at a time, and repeat the details you want to keep. Even then, things can drift during repeated edits, so restate what matters and look closely at each result.
Think of it as a conversation. First: build the scene. Then: keep the scene, move the camera closer. Then: keep everything else, make the machine darker. Then: preserve the composition, add a faint glow inside the machine.
You’re sculpting. You’re not rolling the dice again.
refining an AI image without starting over
References can say what words can’t
Sometimes you know exactly what something should look like, but describing it would take a page. If your tool supports reference images, use them. But don’t just upload one and say “make something like this.” Tell the model what the reference is for:
Use Image 1 for the architecture and proportions. Use Image 2 only for the colour palette. Keep the composition from Image 1 and do not copy its objects.
A single reference can carry structure, proportions, colour, material, pose, composition or atmosphere. You decide which part it contributes. OpenAI’s guide recommends exactly this: identify each input by number and purpose, and explain how they should combine. How you label images depends on the tool, so check how yours refers to them.
reference-image workflows in the Image Prompt Vault
Sometimes the most accurate description is a picture.
The trick is telling the machine which part of the picture you actually mean.
Don’t confuse length with control
This is probably the biggest myth around AI image generation. People find a 700-word prompt online and assume it must be better because it’s longer. Somewhere out there is a 700-word prompt that says “ultra-detailed” eleven times and never mentions where the camera is.
A short prompt can be very precise. A long one can contradict itself. OpenAI’s ChatGPT image guide says one to three clear sentences are enough in most cases, while its API guide suggests labelled sections (scene, subject, details, constraints) once a request gets complicated.
So don’t ask how long a prompt should be. Ask how many visual decisions actually need your attention. If it’s four, write four. If it’s twenty, write twenty.
You’re not trying to impress the model. It doesn’t care how sophisticated your sentences sound. If “a lonely astronaut standing beside a broken rover beneath a huge red planet” gives you the picture you wanted, congratulations, you don’t need a 400-word manifesto. The right prompt is the shortest one that gives you the control you need. Sometimes that’s one sentence. Sometimes it’s a page. The image decides.
A better way to build a prompt
Instead of memorising a rigid formula, try building the prompt in passes. (Google’s guidance for Gemini suggests splitting a complex scene into steps too, and some tools will let you do it as an actual conversation.)
Pass one: the idea. Write the simplest possible description.
An enormous machine beneath an underground city.
Pass two: the world. Where does it exist?
An enormous machine beneath an abandoned underground city, surrounded by old concrete tunnels and collapsed infrastructure.
Pass three: the moment. What’s happening?
The machine has just activated after decades of silence, with faint internal lights beginning to travel through its mechanical structure.
Pass four: the viewer. Where are we?
The viewer stands far back in the tunnel, with a small human silhouette in the foreground.
Pass five: the atmosphere. What does it look like?
Cold humid air, thin ground fog, deep shadows and restrained violet illumination against dark concrete.
Now combine them:
An enormous machine beneath an abandoned underground city, surrounded by old concrete tunnels and collapsed infrastructure. The machine has just activated after decades of silence, faint internal lights travelling through its mechanical structure. The viewer stands far back in the tunnel, with a small human silhouette in the foreground. Cold humid air, thin ground fog, deep shadows and restrained violet illumination against dark concrete. Photorealistic cinematic environment design, physically believable materials, strong sense of scale, no text or logos.
That prompt doesn’t work because it’s long, and it doesn’t work because it follows a magic formula. It works because it describes a specific visual event: what the scene is, what’s happening, where the viewer stands, what sets the scale, what the light is doing, and what should stay out of the way.
The first generation is reconnaissance
Your first image doesn’t have to be perfect. It can be information.
Maybe the machine looks right but is too small. Maybe the colours are perfect but the background is too busy. Maybe the composition works but the machine looks decorative instead of functional. Now you know what the next prompt has to fix.
Instead of “make it better,” try:
Keep the current machine, camera position and colour palette. Increase the machine’s scale by roughly one third so that it extends beyond the upper edge of the frame. Reduce the background structures so the machine remains visually dominant.
Or say the subject, architecture, materials, colours and lighting are all right and only the camera is wrong. Don’t touch the rest. Say so:
Keep the subject, architecture, materials, colours and lighting unchanged. Move the camera to a lower position and tilt it upward slightly so the structure feels more imposing.
You’ve isolated the problem, which makes the next result much more stable. The machine gives you feedback. You respond. That loop is where the craft actually lives.
Different models, different behaviour
There’s no single prompt syntax that behaves identically across every image generator. A prompt that works beautifully in one model may wobble in another.
That doesn’t make the fundamentals useless. It means you should separate visual thinking from model-specific behaviour. Learn to communicate the image first. Then learn how your particular generator reads those instructions.
Google’s Gemini guidance, for example, asks for a described scene rather than a pile of keywords. OpenAI’s guide talks about labelled sections and preserving details across edits. Adobe leans on Firefly’s own reference and style controls. The tools change. The underlying skill doesn’t.
image-generation prompts in the Prompt Vault
The strange part
Eventually, something changes.
You stop thinking of prompts as instructions. You start thinking of them as visual coordinates. You write:
narrow corridor, impossible depth, wet black floor, no visible ceiling, distant warm light, human figure almost lost in scale
and you aren’t really writing prose anymore. You’re describing a place.
The machine interprets it. Sometimes it gets close. Sometimes it opens a door you didn’t know was there.
That’s the interesting part.
A prompt isn’t the picture.
It’s the coordinates you give the machine before you disappear into it.
A checklist before you press Generate
-
Do I know what the image is actually about? If not, simplify the idea.
-
Is the important subject obvious? If not, give it more visual priority. What should the eye notice first, second and third?
-
Does the environment support the subject? If not, it may feel like objects floating in a void.
-
Do I know where the viewer is? If not, add a camera position or framing.
-
Is the light doing something? If not, decide whether it should.
-
Does the colour palette have a purpose? If not, don’t force one.
-
Have I shown the mood instead of only naming it? Swap some abstract adjectives for things the viewer can see.
-
Is every word helping? Delete the decorative ones.
-
What must not change, and what am I happy to leave to the model? Say the first out loud. Let the second go.
Good prompting isn’t about controlling everything. It’s about controlling the right things.
The goal isn’t a perfect prompt
There’s no universal perfect AI image prompt. There’s only a prompt that gives the machine enough direction to produce something worth looking at.
Then you look at it. You notice what went wrong. You tell it what to change. You look again.
Eventually the distance between what you imagined and what appears on the screen gets smaller. That’s prompting. Not keyword collecting, not prompt worship, not writing the longest instruction possible.
Seeing clearly enough to tell a machine what you mean.
And once you can do that, the machine becomes considerably more interesting.
Further reading
The official documentation from the major platforms is worth reading alongside your own experiments:
-
OpenAI’s image prompting guide (developers.openai.com/api/docs/guides/image-prompting) covers descriptive prompts, composition, constraints, reference images and one-change-at-a-time editing.
-
Google’s Gemini image generation documentation (ai.google.dev/gemini-api/docs/image-generation) has prompting examples and strategies, including scene-first descriptions.
-
Adobe’s guide to writing effective text prompts for Firefly (helpx.adobe.com/firefly/web/work-with-images/generate-images/writing-effective-text-prompts.html) covers descriptive prompting and the controls available in Firefly.
The best way to learn is still the least glamorous one: write something, generate it, look at what happened, and try again.