Beyond Words: A Preliminary Study for Multimodal Hate Speech Detection

Abstract

Detecting hate speech is especially challenging when harmful meaning is implicit, contextual, sarcastic, or expressed through a combination of text and images. This preliminary study investigates multimodal hate-speech detection for content such as memes. It explores representations for textual and visual information as well as strategies for combining both modalities in a single detector. Experiments on publicly available datasets show encouraging performance compared with more complex approaches while also exposing the difficulty of interpreting multimodal signals reliably. The work provides an initial basis for studying how visual and linguistic evidence can be jointly modeled to identify harmful content that may not be recognizable from either modality alone.

Publication
_Proceedings of the L Latin American Computer Conference (CLEI 2024), Bahía Blanca, Argentina, https://doi.org/10.1109/CLEI64178.2024.10700167_