Moderation in Generative Services: How It Works
Checks on input and output, why they err in both directions, and what to do about a false positive.
Moderation in generative services is more layered than it appears: it operates at several levels on different principles. Understanding the mechanics explains why a harmless request is sometimes refused while a dubious one goes through.
Levels of checking
- Prompt pre-checks: analysing the text before generation runs. The cheapest level and the most prone to false positives.
- Input image checks: screening uploaded photographs against the rules.
- Output checks: analysing the generated image before showing it to the user.
- Response to reports: reviewing published material following complaints.
Different services use different combinations: some rely mainly on pre-checks, others predominantly on output review.
Why pre-checks err
They work on text without knowing intent. A word that appears in problematic requests also appears in harmless ones: a medical context, a historical description, an analysis of a film scene, an artistic metaphor.
The result is false positives on normal requests and misses where a problematic request happened to avoid the characteristic words.
Why output checks are more reliable
They look at what was actually produced rather than at intent. That removes many false positives but costs more: the generation has already run and the computation is spent.
A practical consequence: in services with output checks a refusal can occur after generation, and the charging policy in that case is worth clarifying.
What is usually closed off
- Sexualised content.
- Images of minors in inappropriate contexts.
- Scenes of violence and cruelty.
- Hateful material.
- Images of real people in invented or compromising situations.
- Content facilitating deception: forged documents, false evidence.
What to do when refused
- Read the stated reason if there is one: it is often specific and tells you what to change.
- Remove words open to a second reading.
- Add context clarifying the purpose: "medical illustration", "historical reconstruction".
- Split a complex request: sometimes the combination triggers rather than individual elements.
- If the refusal repeats, accept it.
What not to do
Attempts to bypass checks through word substitution, transliteration, splitting or encoding. Such actions are detected and usually lead to restricted access — and the act of circumvention is itself read as the intent the rule exists to prevent.
In practice: if a request requires circumvention, it is worth asking whether it is genuinely harmless.
The balance
Moderation is always a compromise: strict settings create false positives and irritate honest users; lenient ones let harmful content through. There is no ideal calibration, and services sit at different points on that spectrum.
The practical conclusion for a user is simple: if you regularly hit false positives on normal tasks, the particular service may be tuned too strictly for your field — and it is worth finding another rather than fighting the filter.
Reports and review
Alongside automated checks there is a response to complaints. If a published image breaches rules or rights, the reporting mechanism is the standard route and usually works faster than trying to make contact directly.
Frequently asked
Why is a harmless request refused?
Pre-checks work on text without knowing intent: a word from problematic requests also occurs in medical, historical and artistic contexts.
What should I do about a false positive?
Rephrase, removing ambiguity and adding context about purpose. If the refusal repeats on a clearly harmless formulation, the words are not the issue.
Can filters be bypassed?
Attempts are detected and usually lead to restricted access, because the act of circumvention itself reads as the intent the rule exists to prevent.
- #модерация
- #как это работает
- #правила