How to Judge a Generative Model: What to Look at When Comparing
Pretty demos say nothing about how a model behaves on your task. Five checks that give a meaningful answer within a few runs.
When choosing between models it is easy to rely on demo galleries. They are useless by definition: they contain the best results from many attempts, often with manual finishing. Meaningful comparison requires your own runs and a clear idea of what to check.
Check one: the hard zones
The first and most informative test is what models traditionally do badly. Request a portrait with hands in frame, lettering on an object, paired jewellery. A model that handles these will handle easy tasks too.
Look not at one good frame but at the proportion of good frames out of five runs. A single hit can be luck.
Check two: instruction following
Write a request with five specific requirements: a particular angle, a clothing colour, a background type, a light source, a facial expression. Count how many were met.
This matters more than artistic quality: a model that makes beautiful pictures of the wrong thing will need three times as many runs on any concrete task.
Check three: robustness under difficult light
Request a scene with non-trivial lighting: rim light, two coloured sources, reflections. This is where models diverge the most.
Check coherence: do shadows fall on one side, do highlights match the sources, does skin tone break into patches.
Check four: likeness in photo mode
If your task is portraits, this is the decisive test. Take one source and run it through the candidate models with an identical request.
Judge by three anchors: eye spacing relative to face width, nose shape, jawline. A general impression of "similar" is deceptive — it survives noticeable divergence of features.
Check five: spread
Run one request five times and look at the spread of quality. A model with a stable average is more practical than one that occasionally produces a masterpiece and occasionally junk.
Practical value is set by the average, because you pay for every run rather than for the best of them.
What not to do when comparing
- Compare on a single frame: spread swamps the difference between models.
- Use different requests for different models: the results are not comparable.
- Judge by galleries and showcases: they are curated.
- Judge by speed: seconds of difference are negligible next to differences in how many runs you need.
What actually matters in practice
After a few comparisons it becomes clear that the decisive property is not artistic quality but predictability. A model that does exactly what was asked saves more than a model with a prettier ceiling.
The second factor is behaviour in your particular niche. A model excellent at landscapes can be mediocre at portraits, and vice versa. There is no universal leader.
How to record results
A comparison is only useful if it can be rechecked. Keep the requests, the sources and every run, including the failures. In a month, when a new version appears, you will have a baseline for comparison — and you will see real progress rather than an impression from fresh demos.
Frequently asked
Can I compare models by the galleries on their sites?
No. Those are the best results from many attempts, often manually finished. Comparison requires your own runs on identical requests.
How many runs make a fair comparison?
At least five per request: the quality spread within one model is often larger than the difference between models on a single frame.
What matters more — beauty or predictability?
Predictability. A model that does exactly what was asked saves more runs than a model with a more impressive ceiling.
- #сравнение моделей
- #качество
- #метод