This paper explores how large language models can support people reviewing harmful online content, including hate speech, harassment, violent extremism and election misinformation.
The study evaluates five ways to combine model output with human judgment, such as filtering clearly non-violative material, identifying possible review errors and providing useful context. Experiments on 50,000 comments and a pilot in a real review queue demonstrate opportunities to improve both reviewer capacity and detection quality. The focus is on designing useful assistance around human review.