Content moderation guidelines that two reviewers apply the same way

Public rules and internal guidelines are two documents

Your community guidelines tell users what is not allowed, in friendly language. Your moderation guidelines tell reviewers exactly how to decide, including the cases you would never publish. Keep them separate, and keep them in sync: every public rule maps to one or more internal categories, and every internal category points back to the public rule a user can be shown in a statement of reasons.

Start from categories, not from words

Lists of banned words age in weeks. Categories last. A practical starting set for most platforms looks like this.

  • Illegal content, with sub categories per market
  • Child safety, always escalated
  • Violence and threats, including self harm
  • Hate and harassment
  • Sexual content and nudity
  • Fraud, scams and spam
  • Impersonation and fake accounts
  • Prohibited goods and services
  • Personal data and doxxing
  • Off platform contact, where your model forbids it

Give every category severity levels

A category without severity forces reviewers to choose between deleting everything and allowing everything. Use three or four levels, and tie each one to an action.

SeverityTypical action
CriticalRemove, suspend the account, escalate immediately
HighRemove and warn or suspend
MediumRemove or restrict reach, warn the user
LowAllow with a label, or no action

Write the line, then write both sides of it

For every category, describe what crosses the line and what stays just inside it. The second half matters most. "Insults directed at another user" is a rule. "Quoting an insult in order to report it" and "banter between friends who both use the word" are the examples that make reviewers agree.

Add examples for every grey area

Examples are the part of the guideline reviewers actually use. Write them as short cases with a decision and a reason. Add a new example every time QA finds two reviewers disagreeing. After a few months, your example library is the most valuable moderation asset you own.

Define escalation triggers before you need them

Some cases are not decided by the reviewer at all. Write down which ones, who receives them, how fast and through which channel.

  • Credible threats to life or of violence
  • Any signal that a minor is at risk
  • Requests from law enforcement or courts
  • High profile accounts and press sensitive cases
  • Content the reviewer cannot classify after reading the guideline

Make it market aware

The same post can be legal in one country and illegal in another. Mark the rules that differ by market and give reviewers the market of the content and the user, so they apply the right version.

Version every change

Every change gets a version number, a date and a short note on what changed and why. Decisions in your audit log reference the version, so you can answer the question "why was this removed in March" a year later.

Test it before launch

Before a guideline goes live, give the same set of real cases to several reviewers separately and compare their decisions. Every disagreement is a sentence to rewrite or an example to add. Repeat until agreement is stable.

Where to go next

If you want help turning your community guidelines into a decision tree, our content moderation process starts exactly there. You can also size the team that will apply it with the plan builder.

Get your moderation plan in two minutes

Tell us your platform type, content, volume, languages and SLA. The plan shows your team per shift, the AI and human split and a monthly estimate.

Get my moderation plan