Skip to main content
Supporting Human Raters with the Detection of Harmful Content Using Large Language Models
  1. publications
  2. ai

Supporting Human Raters with the Detection of Harmful Content Using Large Language Models

Available Media Publication (PDF)
Conference IEEE Symposium on Security and Privacy (S&P) - 2025
Authors Kurt Thomas , Patrick Gage Kelley , David Tao ,
Citation BibTeX
BibTeX
@inproceedings{Thomas2025Supporting,
  title = {Supporting Human Raters with the Detection of Harmful Content Using Large Language Models},
  author = {Kurt Thomas and Patrick Gage Kelley and David Tao and Sarah Meiklejohn and Owen Vallis and Shunwen Tan and Blaž Bratanič and Felipe Tiengo Ferreira and Vijay Kumar Eranti and Elie Bursztein},
  booktitle = {IEEE Symposium on Security and Privacy},
  year = {2025},
  organization = {IEEE}
}

This paper explores how large language models can support people reviewing harmful online content, including hate speech, harassment, violent extremism and election misinformation.

The study evaluates five ways to combine model output with human judgment, such as filtering clearly non-violative material, identifying possible review errors and providing useful context. Experiments on 50,000 comments and a pilot in a real review queue demonstrate opportunities to improve both reviewer capacity and detection quality. The focus is on designing useful assistance around human review.

newsletter signup
newsletter signup