Where AI is strong
- Volume: models score every item in milliseconds, at any hour
- Clear cases: explicit nudity, known spam, duplicates and banned links
- Consistency on the obvious: the same input gets the same score
- Triage: ranking a queue so people see the riskiest items first
- Protection: catching the worst media before a person ever sees it
Where AI fails
- Context: a slur quoted in a report looks like a slur
- Intent: sarcasm, satire and reclaimed language
- Patterns across messages: a romance scam is polite one message at a time
- Novelty: new slang, coded emoji and fresh evasion tricks
- Local norms and law: what is acceptable differs by market
- Appeals: the user adds context the model never saw
Where human moderation fails
People are not perfect either. They get tired, they drift from the policy over weeks, they disagree with each other on grey areas, and they are harmed by too much exposure to graphic content. Good human moderation is designed around those weaknesses: clear guidelines with examples, weekly QA and calibration, rotation and exposure limits, and AI in front so people see less of the worst material.
The split that works
Score everything with AI. Set two thresholds per category. Above the upper one, the item closes automatically. Below the lower one, it passes. Everything between goes to a trained person. Then check both sides: QA samples human decisions and automatic decisions every week.
Moving the thresholds is a business decision
Tighter thresholds send more items to people: higher cost, fewer automatic mistakes. Looser thresholds do the opposite. There is no correct setting, only a trade off you should make with data from your own queues, and revisit when abuse patterns change.
What this means for cost
In most content mixes, AI resolves a large share of text and image items and a smaller share of profiles, video and appeals. Appeals should always go to a person. The human team you need is sized on what remains, and that is where the plan builder starts.
Read more on AI content moderation and how our moderation software handles thresholds.