How an AI Photo Generator Works: From Text to Image
What happens between sending a request and seeing a picture, why the model removes noise rather than drawing, and what follows from that in practice.
Understanding the mechanics is not required to use a generator, but it sharply reduces the number of pointless attempts. Most "strange" model behaviour stops being strange once it is clear what actually happens inside.
The model does not draw — it removes noise
Intuitively it feels as though a generator draws: outline first, then detail, the way a person would. The process is in fact the reverse. The model starts from random noise — a screen of meaningless dots — and step by step removes whatever does not match the request.
At each step it answers one question: "if this image really contains what the request describes, which noise here is superfluous?" Having removed some, it repeats. After a few dozen steps an image emerges from the chaos.
Where understanding of the request comes from
The text is converted into a set of numbers representing meaning — not individual words but their combined sense. Those numbers set the direction in which noise is removed. An important consequence follows: the model responds to the overall meaning of a phrase, not to isolated keywords.
So "a portrait of a woman in a red dress" and "woman, red dress, portrait" give similar but not identical results: the meaning differs slightly, and so does the direction.
Why the same request gives different pictures
The starting noise is random. The same request beginning from different noise arrives at different images — like two sculptors working from one sketch on different blocks of marble.
This is a property rather than a flaw. It is what lets you obtain variants and choose, instead of rewriting the request for the sake of variety.
Why hands and text are hard
Noise removal is local: each region of the image is refined largely independently. For a face that works — it is a coherent whole seen many times in the data. For a hand it works worse: fingers are repeating similar objects, and the model easily loses count while refining each one separately.
Text is harder still: letters must form a specific sequence, while the model works on visual plausibility. It draws something that looks like text, because looking right is what it optimises.
What resolution means
Models are trained on images of a particular size and work confidently near it. A request for a very large image is usually handled differently: generation at the native size, then enlargement by a separate model that paints in detail.
Hence a practical observation: detail in an upscaled image did not "emerge" — it was invented during enlargement. For a portrait that means the skin texture after upscaling is not yours.
Generating from a photo
When an image is supplied, the process changes: instead of pure noise, the model starts from your photo, noised to a certain degree. The heavier the noising, the freer the result and the less of the source survives.
This explains a trade-off familiar from practice: strong stylisation inevitably reduces likeness, because stylisation needs freedom, and freedom is bought by erasing source information.
Why the model "ignores" part of the request
Direction is set by the whole request at once, and elements compete. A long description with a dozen requirements blurs the direction: each requirement gets less weight. A short, clear request is followed more precisely simply because it contains fewer competing signals.
For the same reason negations work badly: "without a hat" introduces the concept of a hat, and the model often draws one. Direction is set by the presence of meaning, not its absence.
What this means for your work
- State what should be present, not what should be absent.
- Keep the request short and prioritised: the essentials first.
- Get variants by rerunning, not by rewriting the request.
- Do not expect precision in small repeating detail — hands, text, patterns.
- Remember that stylisation and likeness pull in opposite directions.
Frequently asked
Why does the same request give different pictures?
Generation starts from random noise. Different starting noise with one request leads to different images — a property of the method, not a fault.
Why does the model draw unreadable text?
It optimises visual plausibility, not letter sequences. You get something that looks like text, because looking right is precisely its objective.
Why does stylisation reduce likeness?
Stylisation needs more freedom, and freedom comes from noising the source more heavily. The more freedom, the less original information about appearance survives.
- #как это работает
- #теория
- #диффузия