# The anatomy of a photo-realistic image prompt

Canonical: https://romeoapps.app/en/larpgpt/anatomy-of-a-photo-prompt/
Author: Roméo Gambino · Published: 2026-09-06 · App: LarpGPT (https://romeoapps.app/en/larpgpt/)

What each part of a prompt does inside an image generator, why one sentence per decision works better than a wall of adjectives, and how to adapt a prompt without breaking the thing that made it work.

## A prompt is a brief, not a spell

The most persistent misconception about image prompts is that certain words are magic and the job is to collect them. What actually works is closer to briefing a photographer who cannot ask questions. You are describing a scene that someone has to set up, light and shoot, and every decision you leave unstated is a decision the model makes for you, usually toward the most average version of what you asked.

## One sentence per decision

Long prompts fail not because they are long but because they are undifferentiated. Twenty adjectives in one sentence compete with each other and the model averages them. The same content split into short sentences, each carrying one decision, produces a far more predictable result. It also makes the prompt editable, because you can see which sentence to change when something is wrong.

## The subject comes first and carries the most weight

Position matters. What appears early in a prompt is generally weighted more heavily, so the subject belongs at the front, described concretely. A concrete subject is one you could photograph: an object, a place, a situation with physical properties. Abstract subjects give the model nothing to build from and it will invent something generic to fill the space.

## Describe what would be visible, not what is felt

This is the single most useful discipline. A word like melancholic is an interpretation, and the model has to guess which visual arrangement produces it. Rain on a window, a single lit lamp, an empty chair: those are things a camera records. Write the causes rather than the effect, and the effect arrives on its own with far more consistency.

## The scene sets the constraints

The setting is not decoration. It determines what light is available, what surfaces exist to reflect it, what can appear in the background, and what scale everything reads at. A subject described in detail with no stated setting will be placed somewhere generic, and the generic setting will then contradict half the choices you made about the subject.

## The moment narrows a scene to a photograph

Time of day, weather and season do more work than most people expect, because each one carries a whole lighting setup with it. Late afternoon in autumn implies a low warm sun, long shadows and a particular colour cast, all from four words. Naming the moment is the most efficient sentence in most prompts.

## Light is a decision, not an afterthought

Photographs are largely made of light, and an unstated lighting choice is the fastest route to a flat, evenly-lit image that looks synthetic. Say where the light comes from and what kind it is. Window light from the left. A single overhead source. Overcast sky. Direct sun through leaves. Each produces a different image from the same subject and scene.

## Hard and soft light do different jobs

Light from a small or distant source is hard: sharp shadow edges, high contrast, visible texture. Light from a large or diffused source is soft: gradual shadows, low contrast, gentler surfaces. Naming which one you want tells the model more about the mood of an image than any adjective about mood does.

## Composition instructions the model can follow

Framing terms are useful because they are unambiguous. Close-up, wide shot, eye level, low angle, overhead. Whether the subject sits centred or off to one side. What is in front of it and what is behind. These are the words a photographer would use on set, and they translate cleanly because they describe geometry rather than taste.

## Lens language, and what it actually does

A wide focal length exaggerates depth and makes near things large. A long focal length compresses distance and flattens layers. A wide aperture throws the background out of focus and isolates the subject. These terms work in prompts because they correspond to real optical behaviour, and asking for a shallow depth of field is more precise than asking for a blurry background.

## Colour, restrained

Naming two or three colours produces a coherent image. Naming eight produces mud, because the model tries to include all of them and none dominates. If a colour matters, attach it to an object rather than to the whole picture: a red door reads better than a red-tinted scene, and it leaves the rest of the palette free to follow the light.

## Change one thing at a time

When a result is close but wrong, the instinct is to rewrite the prompt. Resist it. Change one sentence, regenerate, and see what moved. This is slower per attempt and much faster overall, because you learn what each sentence is doing. A rewritten prompt that works teaches you nothing you can reuse tomorrow.

## Keep the versions you discard

A prompt that produced something interesting but not what you wanted is worth keeping. Prompt libraries are built from those, not from the successes, because the near-misses are what you adapt when the next brief is slightly different. A note on what the result actually looked like is worth more than the prompt alone.

## Adapt without breaking the structure

When you reuse a prompt, swap the subject and the setting and leave the light, composition and lens sentences alone at first. Those are the parts that make an image look photographic, and they are the parts people delete when a prompt feels too long. Change them deliberately, one at a time, once the new subject is working.

## Say what the image is

Images made this way are not photographs, and the standards for saying so already exist. The IPTC Digital Source Type vocabulary provides values that distinguish a digital capture sampled from real life from media created using generative AI, and the C2PA specification defines how provenance information travels with a file. Where the image might be taken for a record of something real, label it. LarpGPT provides prompts and reference material, and what you generate is yours to use and yours to disclose.

## Sources

- IPTC Digital Source Type NewsCodes — https://cv.iptc.org/newscodes/digitalsourcetype/
- C2PA technical specification — https://c2pa.org/specifications/specifications/2.1/index.html
- IPTC Photo Metadata Standard — https://www.iptc.org/std/photometadata/specification/IPTC-PhotoMetadata

