Inference-Time Neurosymbolic LLM Integration for Content Moderation
Open Access DepositedLarge Language Models (LLMs) have been used widely in natural language processing (NLP) applications due to their strong contextual understanding abilities. Content moderation is also an area where LLMs are being explored for automating toxicity detection in content moderation due to their strong contextual inference and natural language understanding abilities. Even with their exemplary abilities in natural language processing, these LLMs have not been able to provide the desired levels of accuracy in toxicity detection. In addition, the black-box nature or the opaqueness of the decision-making process is a key drawback in LLMs that leads to a lack of trust in the decisions made by these models. As a result, social media companies continue to rely excessively on human moderators. Several approaches to improve the performance of LLMs have been explored. These approaches include zero-shot prompt engineering and multi-shot in-context learning. Despite these approaches, LLMs have only shown marginal improvement in the toxicity classification accuracy, continuing to render these very powerful models as untrustworthy for automating toxicity classification in content moderation. This praxis identifies the key reason for the poor performance of these LLMs in toxicity detection as the lack of content moderation domain-specific knowledge. These LLMs are built as general-purpose, natural language models and they lack the "awareness" of what constitutes toxic and proposes a hybrid neurosymbolic approach for toxicity classification in content moderation to address these issues. This hybrid approach combines the strengths of two separate modeling paradigms, the symbolic-only model and the neural-only model where the symbolic-only model provides structured knowledge and rules while the neural-only model (an LLM) provides the contextual inference capabilities. The performance of this neurosymbolic model is tested using the real-world dataset from the benchmark study by Roy et al. (2023). The results of the Neurosymbolic model are benchmarked against the against the results of the neural-only models from the benchmark study, in two configurations i.e. zero-shot vanilla-prompting and multiple prompt-enhanced variations. The results of the Neurosymbolic model are also tested against the results of a symbolic-only model that was also developed as part of this praxis. Accuracy and f1-score are used as the primary evaluation metrics. The results indicate that the neurosymbolic model consistently outperforms the neural-only baseline in both the vanilla-prompt and prompt-enhanced configurations. It also outperformed the symbolic-only baseline across both metrics. In addition to the improved performance, the model was also able to generate structured explanations of the moderation decisions. The architecture’s flexible design supports future enhancements and positions it as a promising foundation for more robust, scalable, and transparent content moderation systems.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.
| Thumbnail | Title | Date Uploaded | Visibility | Actions |
|---|---|---|---|---|
|
|
Srinivasa_gwu_0075A_17476.pdf | 2025-12-11 | Open Access |
|