US & Europe edition
Merxtio
Markets
  • S&P 500 ETF769.38 0.22%
  • Bitcoin77,661.00 1.95%
  • Ethereum2,434.69 2.65%
  • Solana103.74 1.21%
  • EUR/USD1.1643
  • GBP/USD1.3583
  • USD/JPY159.68

How Do You Get an AI Image That Matches What You Asked For?

Image models do not follow instructions, they continue descriptions. That one difference explains most bad output — and the license, not the prompt, decides whether you can use the result.

By Merxtio Staff

5 min read

A dark monitor displaying a colorful particle render beside a panel of controls
Photo by Egor Komarov on Pexels

You typed "create an image of a mountain at sunset" and got something that looks like a screensaver from 2009. The tool is not broken and you are not bad at this. You gave it an instruction, and it wanted a description.

This page covers the difference, the eight elements that make a prompt reliable, and the thing most image guides leave out entirely — which tool's license actually lets you put the result on a client invoice. About eight minutes.

Describe the picture, do not order one

Chat models follow instructions. Image models continue descriptions. Almost every disappointing result traces back to writing a request when you should have written a caption.

"Create an image of a mountain at sunset" is a sentence about your intentions. "A majestic mountain at sunset, dramatic clouds, warm alpenglow on the rock face" is a sentence about the picture. The second works better because it gives the model far more to continue from, and none of it is spent on words like create and image of that describe nothing visual.

The same applies to a subject. "A dog" is unpredictable. "A golden retriever puppy, soft natural lighting, warm and joyful mood, photorealistic" is usable. That is not more effort — it is three extra clauses that took four seconds.

The eight elements

Not every image needs all eight. Reaching for the relevant ones, in roughly this order, is what separates a reliable prompt from a lucky one.

Prompt architecture, weak version against strong
FeatureWhat it settlesExample
SubjectWho or what is in frame.Confident executive in her mid-forties
SettingWhere they are.Modern glass office, city visible behind
ActionWhat is happening, which stops stiff portrait poses.Reviewing documents at a standing desk
StyleThe single highest-leverage element. Names a visual tradition.Editorial photography
LightingDoes more for realism than any quality keyword.Soft natural window light from the left
MoodThe emotional register, which changes color and composition.Professional, focused
CameraLens and depth language the model has seen in captions.Shallow depth of field, 85mm
ParametersTool-specific switches for ratio and quality.Aspect ratio 16:9

"Businessman in office" produces a lottery ticket. The eight-element version above produces something you could put in a deck, and it takes under a minute once the shape is familiar.

If you add only one thing, add style. A subject with the right style modifier beats a heavily detailed subject with no style context, because the style word activates a whole visual tradition at once. Editorial photography, oil painting, isometric illustration, architectural photography — each one moves the output further than another sentence of description would.

Lighting is the close second, and the one people skip. Golden hour, soft diffused, dramatic side lighting, blue hour will do more for how real an image looks than stacking 8K, hyperrealistic and masterpiece, which are quality signals rather than visual instructions.

Parameters are tool-specific

Midjourney takes switches at the end of a prompt. They are worth knowing if you use it, and they do not transfer to other tools.

  • Aspect ratio — set it before generating rather than cropping after, because composition is built around the frame. Square for feed posts, 16:9 for slides and thumbnails, 9:16 for vertical video.
  • Quality — a low setting for fast exploration, the high one only for finals. Most people burn their allowance rendering drafts at full quality.
  • Stylize — how much of the tool's own aesthetic gets applied. Low keeps it close to your literal description; high produces something prettier and less like what you asked for. When output looks lovely but wrong, this is usually why.

That last one is worth internalizing, because it is the setting that most often explains the gap between what you described and what arrived.

The part most guides skip

Prompt craft decides whether you get the image. Licensing decides whether you can use it.

Entry price against the tier a business actually needsUSD per month
Entry price against the tier a business actually needs. USD per month.
ItemUSD per month
Adobe Firefly, any paid plan$9.99
Midjourney Basic$10
Midjourney Pro, required above $1M revenue$60

Vendor pricing and license terms, checked 21 August 2026. Midjourney tiers run $10, $30, $60 and $120.

The entry prices look identical. The obligations do not.

Commercial terms, checked 21 August 2026
FeatureAdobe FireflyMidjourney
Commercial useIncluded on every paid plan.Permitted, but businesses over $1M annual revenue must be on Pro or Mega.
IP indemnificationFull commercial copyright indemnification — the only major generator offering it.None, at any tier.
Training dataLicensed and public-domain sources.Broad web-scale training.
Best forClient work, advertising, anything with a legal review.Editorial, personal, internal, concept work.

Indemnification is the word that matters. It means the vendor will stand behind you if someone claims the output infringes their work. Firefly offers it; Midjourney explicitly does not. For a personal project that is irrelevant. For a client campaign it is the entire decision, and it is why the better image generator is often the wrong purchase.

Build a prompt that works

  1. Write the caption, not the request

    You’ll have: A description with no instruction verbs in it. · about 2 minutes

    Delete "create", "generate", "make me an image of". They occupy space and describe nothing.

    What is left should read like a caption under a photograph in a magazine. If it reads like an email to a designer, rewrite it.

  2. Add style and lighting before anything else

    You’ll have: A prompt that already produces usable output. · about 2 minutes

    One style word and one lighting phrase. This is the largest single jump in quality available, and it is two clauses.

    Resist adding 8K, masterpiece and award-winning at this stage. They are weak signals compared with naming an actual visual tradition, and they crowd out the words that carry real information.

  3. Fill the gaps that still look wrong

    You’ll have: A prompt reliable enough to reuse. · about 5 minutes

    Now work through the remaining elements against what the output got wrong. Stiff pose means the action element is missing. Wrong emotional register means mood. Flat and snapshot-like means camera language.

    Change one element at a time. Changing three at once teaches you nothing about which one worked.

  4. Save the prompt as a template

    You’ll have: A repeatable house style rather than a lucky result. · about 3 minutes

    Replace the subject with a blank and keep everything else. Style, lighting, mood and camera are what make a set of images look like they belong together, so freezing them is how you get visual consistency without any special feature.

The description-versus-instruction distinction is the same mechanism covered in the prompting guide — both come down to giving the model more of the right thing to continue from. And if you are still deciding whether an image tool earns a place in your stack at all, the four-question test applies here as much as anywhere.

Questions people ask

Why do my AI images not look like what I described?
Usually one of two things. Either the prompt is an instruction rather than a description, so most of its words carry no visual information — or the tool is applying a heavy dose of its own aesthetic. In Midjourney that is the stylize setting, and lowering it brings output closer to your literal description.
How do I write a good Midjourney prompt?
Describe the picture in this order: subject, setting, action, style, lighting, mood, camera, then any parameters. Not all eight are needed every time, but style and lighting are the two that move quality most, and they are the two people most often leave out.
Can I use AI-generated images commercially?
It depends entirely on the tool. Adobe Firefly includes commercial rights on all paid plans. Midjourney permits commercial use but requires businesses over $1 million in annual revenue to be on Pro or Mega, and offers no IP indemnification at any tier. Check the license before the campaign, not after.
Does Midjourney own the images I generate?
No — paid subscribers own their outputs under Midjourney terms. The gap is not ownership, it is indemnification: if a third party claims the image infringes their work, Midjourney does not stand behind you. That is the risk Adobe Firefly is priced to remove.
What is the best AI image generator for business?
Adobe Firefly, on licensing rather than looks. It carries commercial rights on every paid plan from $9.99 a month and full copyright indemnification. Midjourney generally produces better images, which makes it the better tool and the riskier business purchase.
How do I get consistent style across images?
Freeze everything except the subject. Keep the same style, lighting, mood and camera language in every prompt and swap only what is in frame. Consistency comes from the constant half of the prompt, which is why saving it as a template matters more than any single generation.