Generating From Scratch and Generating From a Photo: The Practical Difference
Two modes of working with a model solve different problems. What each gives you, where the boundary runs, and how to choose for a given outcome.
Almost every generation service offers two fundamentally different ways to get an image: describe what you want in words, or upload a photograph as the basis. The difference runs deeper than it appears, and the choice of mode determines what is obtainable at all.
Generating from scratch
You supply only text. The model starts from random noise and is entirely free in what to draw: any appearance, any scene, any style. The single constraint is what you described.
The strengths are obvious: maximum freedom, no requirements on source material, the ability to get something that does not exist. The weakness is also single but substantial: you cannot get a specific person. Even the most detailed description of appearance yields a similar type, not an identity.
Generating from a photo
You supply an image plus a description of what to do with it. The model starts not from noise but from your photograph, and the result inherits its properties — appearance above all.
The strength: recognisability. The weakness: freedom is limited. The more you want to change the scene, the more you must let go of the source, and likeness goes with it.
Where the boundary runs
The practical rule is simple. If a specific person must appear in the result — photo mode, no alternatives. If a specific person is not needed, generation from scratch gives a better result: it spends no capacity holding on to a face and handles the scene better.
- A portrait of yourself in a new style — from a photo.
- An illustration for a post with no recognisable faces — from scratch.
- An avatar colleagues should recognise — from a photo.
- A background image, a cover, an abstract scene — from scratch.
- A character for a project who does not exist — from scratch.
Intermediate modes
Between the two extremes there are variants worth knowing.
- Scene reference. The uploaded photo sets composition, light and setting rather than appearance — while the face comes from another source or description.
- Detail refinement. The source is an already-generated image: the model refines it without changing the composition.
- Combined input: one photo for the scene, several for appearance. This is how scenarios that reproduce someone else's frame work.
What happens to style
From scratch, style is set entirely by the description and the model follows it precisely. From a photo, style competes with the source: if the shot has hard light and you ask for soft studio lighting, the model must rewrite the illumination — and does so worse the wider the gap.
Hence a practical note: in photo mode, choose a style not too far from the source's shooting conditions. A radical change of light is possible but is paid for in likeness.
Material requirements
Generating from scratch requires nothing but wording. Generating from a photo requires a usable shot, and that is not a formality: source quality is the primary factor in the outcome, more so than style choice or prompt length.
The typical mistake
The most common is trying to get a recognisable portrait through a textual description of appearance. "A woman of thirty, dark hair, green eyes" specifies a type, and the model will draw a plausible person — but not you, and not your acquaintance.
The reverse mistake also happens: uploading a photo where none is needed. If the frame should contain no recognisable face, a source only constrains the model and degrades the scene.
Cost and speed
The two modes usually cost similarly per run but differ in how many runs it takes to arrive. From scratch tends to need iteration on wording; from a photo, selection of source and style. In both cases the greatest saving is made in preparation, not in process.
Frequently asked
Can I get a specific person from a text description?
No. A description of appearance specifies a type, not an identity. A recognisable person requires photo-based generation.
Why does heavy stylisation lose the likeness?
Stylisation needs freedom, and freedom comes from loosening the tie to the source. The further the style is from the original shooting conditions, the more appearance drifts.
When does an input photo only get in the way?
When the frame needs no recognisable face. Then the source constrains the model and degrades the scene — generating from scratch is better.
- #режимы
- #сравнение
- #практика