How Do You Get an AI Image That Matches What You Asked For?
Image models do not follow instructions, they continue descriptions. That one difference explains most bad output — and the license, not the prompt, decides whether you can use the result.
By Merxtio Staff

You typed "create an image of a mountain at sunset" and got something that looks like a screensaver from 2009. The tool is not broken and you are not bad at this. You gave it an instruction, and it wanted a description.
This page covers the difference, the eight elements that make a prompt reliable, and the thing most image guides leave out entirely — which tool's license actually lets you put the result on a client invoice. About eight minutes.
Describe the picture, do not order one
Chat models follow instructions. Image models continue descriptions. Almost every disappointing result traces back to writing a request when you should have written a caption.
"Create an image of a mountain at sunset" is a sentence about your intentions. "A majestic mountain at sunset, dramatic clouds, warm alpenglow on the rock face" is a sentence about the picture. The second works better because it gives the model far more to continue from, and none of it is spent on words like create and image of that describe nothing visual.
The same applies to a subject. "A dog" is unpredictable. "A golden retriever puppy, soft natural lighting, warm and joyful mood, photorealistic" is usable. That is not more effort — it is three extra clauses that took four seconds.
The eight elements
Not every image needs all eight. Reaching for the relevant ones, in roughly this order, is what separates a reliable prompt from a lucky one.
| Feature | What it settles | Example |
|---|---|---|
| Subject | Who or what is in frame. | Confident executive in her mid-forties |
| Setting | Where they are. | Modern glass office, city visible behind |
| Action | What is happening, which stops stiff portrait poses. | Reviewing documents at a standing desk |
| Style | The single highest-leverage element. Names a visual tradition. | Editorial photography |
| Lighting | Does more for realism than any quality keyword. | Soft natural window light from the left |
| Mood | The emotional register, which changes color and composition. | Professional, focused |
| Camera | Lens and depth language the model has seen in captions. | Shallow depth of field, 85mm |
| Parameters | Tool-specific switches for ratio and quality. | Aspect ratio 16:9 |
"Businessman in office" produces a lottery ticket. The eight-element version above produces something you could put in a deck, and it takes under a minute once the shape is familiar.
If you add only one thing, add style. A subject with the right style modifier beats a heavily detailed subject with no style context, because the style word activates a whole visual tradition at once. Editorial photography, oil painting, isometric illustration, architectural photography — each one moves the output further than another sentence of description would.
Lighting is the close second, and the one people skip. Golden hour, soft diffused, dramatic side lighting, blue hour will do more for how real an image looks than stacking 8K, hyperrealistic and masterpiece, which are quality signals rather than visual instructions.
Parameters are tool-specific
Midjourney takes switches at the end of a prompt. They are worth knowing if you use it, and they do not transfer to other tools.
- Aspect ratio — set it before generating rather than cropping after, because composition is built around the frame. Square for feed posts, 16:9 for slides and thumbnails, 9:16 for vertical video.
- Quality — a low setting for fast exploration, the high one only for finals. Most people burn their allowance rendering drafts at full quality.
- Stylize — how much of the tool's own aesthetic gets applied. Low keeps it close to your literal description; high produces something prettier and less like what you asked for. When output looks lovely but wrong, this is usually why.
That last one is worth internalizing, because it is the setting that most often explains the gap between what you described and what arrived.
The part most guides skip
Prompt craft decides whether you get the image. Licensing decides whether you can use it.
| Item | USD per month |
|---|---|
| Adobe Firefly, any paid plan | $9.99 |
| Midjourney Basic | $10 |
| Midjourney Pro, required above $1M revenue | $60 |
The entry prices look identical. The obligations do not.
| Feature | Adobe Firefly | Midjourney |
|---|---|---|
| Commercial use | Included on every paid plan. | Permitted, but businesses over $1M annual revenue must be on Pro or Mega. |
| IP indemnification | Full commercial copyright indemnification — the only major generator offering it. | None, at any tier. |
| Training data | Licensed and public-domain sources. | Broad web-scale training. |
| Best for | Client work, advertising, anything with a legal review. | Editorial, personal, internal, concept work. |
Indemnification is the word that matters. It means the vendor will stand behind you if someone claims the output infringes their work. Firefly offers it; Midjourney explicitly does not. For a personal project that is irrelevant. For a client campaign it is the entire decision, and it is why the better image generator is often the wrong purchase.
Build a prompt that works
Write the caption, not the request
You’ll have: A description with no instruction verbs in it. · about 2 minutes
Delete "create", "generate", "make me an image of". They occupy space and describe nothing.
What is left should read like a caption under a photograph in a magazine. If it reads like an email to a designer, rewrite it.
Add style and lighting before anything else
You’ll have: A prompt that already produces usable output. · about 2 minutes
One style word and one lighting phrase. This is the largest single jump in quality available, and it is two clauses.
Resist adding 8K, masterpiece and award-winning at this stage. They are weak signals compared with naming an actual visual tradition, and they crowd out the words that carry real information.
Fill the gaps that still look wrong
You’ll have: A prompt reliable enough to reuse. · about 5 minutes
Now work through the remaining elements against what the output got wrong. Stiff pose means the action element is missing. Wrong emotional register means mood. Flat and snapshot-like means camera language.
Change one element at a time. Changing three at once teaches you nothing about which one worked.
Save the prompt as a template
You’ll have: A repeatable house style rather than a lucky result. · about 3 minutes
Replace the subject with a blank and keep everything else. Style, lighting, mood and camera are what make a set of images look like they belong together, so freezing them is how you get visual consistency without any special feature.
The description-versus-instruction distinction is the same mechanism covered in the prompting guide — both come down to giving the model more of the right thing to continue from. And if you are still deciding whether an image tool earns a place in your stack at all, the four-question test applies here as much as anywhere.
Questions people ask
- Why do my AI images not look like what I described?
- Usually one of two things. Either the prompt is an instruction rather than a description, so most of its words carry no visual information — or the tool is applying a heavy dose of its own aesthetic. In Midjourney that is the stylize setting, and lowering it brings output closer to your literal description.
- How do I write a good Midjourney prompt?
- Describe the picture in this order: subject, setting, action, style, lighting, mood, camera, then any parameters. Not all eight are needed every time, but style and lighting are the two that move quality most, and they are the two people most often leave out.
- Can I use AI-generated images commercially?
- It depends entirely on the tool. Adobe Firefly includes commercial rights on all paid plans. Midjourney permits commercial use but requires businesses over $1 million in annual revenue to be on Pro or Mega, and offers no IP indemnification at any tier. Check the license before the campaign, not after.
- Does Midjourney own the images I generate?
- No — paid subscribers own their outputs under Midjourney terms. The gap is not ownership, it is indemnification: if a third party claims the image infringes their work, Midjourney does not stand behind you. That is the risk Adobe Firefly is priced to remove.
- What is the best AI image generator for business?
- Adobe Firefly, on licensing rather than looks. It carries commercial rights on every paid plan from $9.99 a month and full copyright indemnification. Midjourney generally produces better images, which makes it the better tool and the riskier business purchase.
- How do I get consistent style across images?
- Freeze everything except the subject. Keep the same style, lighting, mood and camera language in every prompt and swap only what is in frame. Consistency comes from the constant half of the prompt, which is why saving it as a template matters more than any single generation.