Why AI Struggles With Hands, Text and Symmetry

7 min

Three classic weak spots of generation share one cause. What model developers do about it, and what you can do yourself when composing a frame.

Six fingers on a hand became the calling card of generated images. Together with unreadable text on signage and mismatched earrings, it forms the trio of the most recognisable defects. All three share one cause, and understanding it helps you route around the problem.

The shared cause: locality

A model processes an image as a set of interrelated regions, but the link between distant regions is weaker than between neighbours. Each fragment is refined largely on its own — which makes generation fast and allows work at large sizes.

The price is difficulty with anything requiring a global account: how many fingers there are in total, which letter comes third, whether the earrings match. Locally each finger looks right; the problem appears only when you look at the whole hand.

Why hands specifically

A hand is a uniquely awkward object. It consists of repeating similar elements that can occupy an enormous number of mutual positions, frequently occlude one another, and change apparent shape with any rotation.

Add the statistics of training data: in photographs, hands are most often partly hidden, blurred by motion, or small. The model has seen many examples, but most of them are not crisp anatomical references.

Why text

An image of text must satisfy a requirement alien to visual plausibility: a specific sequence of specific characters. The model optimises how much the picture resembles a real one, and "looks like text" is not the same as "is text".

Modern models handle this noticeably better than earlier ones, especially short words in large type. But long text at small sizes remains a risk zone, and letters on background signage still crumble.

Why symmetry

Paired objects — earrings, buttons, fabric patterns, interior elements — require agreement between regions processed independently. The right earring does not "know" what the left one looks like.

The face is the exception: it is such a frequent and coherent object that the model treats it as a single whole. That is exactly why faces come out better than hands, though they are comparable in complexity.

What to do when composing

  1. Choose a composition without hands in focus: lowered, behind the back, out of frame, in pockets.
  2. Avoid gestures near the face — they combine all three problems at once: anatomy, occlusion, and viewer attention.
  3. Remove text: clothing without lettering, backgrounds without signage.
  4. Take off symmetrical paired jewellery for close framing.
  5. A simple background reduces the number of small details requiring agreement.

What to do with a finished frame

Local defects are often cheaper to crop away than to regenerate. Trimming the bottom of a frame with hands takes a second; chasing correct anatomy through repeated runs takes several attempts with an uncertain outcome.

For paired details, mirroring one half in an editor is sometimes enough. For text — replacing it with a neutral background, or with real type added on top.

Where the technology is heading

The weak spots are gradually closing. Hands come out correct far more often in new models than two or three years ago; short words now work. The reasons are both larger datasets and architectural changes that improve agreement between regions.

But the principle holds: while a model optimises plausibility rather than rules, tasks with hard rules will come harder than tasks with soft criteria. A frame that does not depend on such rules will always be the safer bet.

Frequently asked

Why do faces come out better than hands?

The model treats a face as a coherent object seen an enormous number of times in the data. A hand is repeating elements refined largely independently, and the model loses count of the fingers.

Can I ask the prompt for correct hands?

Wording barely helps: the problem is not understanding the request but how the image is built. Choosing a composition where hands are not in focus is more reliable.

Will it get better over time?

It already has: short lettering and hand anatomy are noticeably more reliable in new models. But tasks with hard rules will remain harder than tasks with soft criteria.

  • #артефакты
  • #ограничения
  • #разбор

Try it in the studio

Upload a photo and pick a style — one step from prompt to result.

Open the studio

Read next

Try it on this topicAI Photo Generator

PromptsAll articles