Who should moderate ‘toxic’ online content: AI, professionals or users?

Online platforms increasingly combine AI and human moderation, but do they make the same decisions? In a new open-access study, Dr. Aviv Barnoy compares how a platform’s AI system, professional moderators and registered users judged the same real online comments. The findings show that these actors apply different thresholds and may therefore be most useful when combined rather than treated as interchangeable alternatives.

Content moderation is often discussed as a choice between humans and AI. In practice, however, many platforms already use hybrid systems in which automated tools, professional moderators and users can all play a role. A new study by Dr. Aviv Barnoy, published in Media and Communication, examines what each of these actors contributes.

Through a collaboration with OpenWeb, Barnoy gained rare access to data connected to an actual moderation environment. The study links three types of judgments about the same real comments: the initial disposition of OpenWeb’s AI-assisted moderation system, final decisions made by professional moderators, and judgments from 863 registered users. Together, the users made 8,588 decisions about whether comments should remain online or be removed.

The study also builds on earlier conceptual work that defines problematic or “toxic” online content through specific norm violations rather than treating toxicity as one broad category. It examines five moderation categories: violence, doxxing, hate speech, sexually explicit content and personal attacks. These categories involve different norms and risks, and may therefore require different moderation thresholds and responses.

The results show substantial differences between the three decision-makers. Users chose to leave roughly two-thirds of the comments online. Their decisions were considerably closer to the AI system’s initial publish-or-hold orientation than to professional moderators’ final publish-or-block decisions. Professional moderators were more restrictive overall. Disagreement between AI and professional moderators was concentrated in particular comments and categories, especially hate speech.

This does not mean that AI is more accurate or more legitimate than professional moderation. Instead, Barnoy argues that the three actors can fulfil different functions. AI can support scalable triage, professional moderators can provide policy- and risk-sensitive adjudication, while users can provide information about community expectations and tolerance thresholds. The findings therefore point toward the potential value of carefully designed hybrid moderation systems rather than searching for one universal moderation tool.

Users’ explanations also reveal an important distinction between recognizing problematic content and deciding that it should be removed. People frequently acknowledged that a comment was objectionable while still judging it insufficiently harmful to justify removal. This suggests that identifying a norm violation and deciding on the appropriate sanction should not automatically be treated as the same question.

Researcher
More information

You can find the full article here.

Compare @count study programme

  • @title

    • Duration: @duration
Compare study programmes