Who Judges the Frame? Auditing Multimodal LLM Judges for News Framing Across Event-Level Perspectives

Abstract

News coverage of major world events is shaped not only by what is reported, but also by how events are framed through text, images and their combination. At the same time, Large Language Models (LLMs), including multimodal LLMs, are increasingly used as scalable instruments for analysing framing, sentiment, ideological slant and perspective differences in multimodal media datasets. This creates a methodological challenge. When used as measurement instruments, LLM outputs may reflect not only content properties, but also model-specific tendencies, prompt design choices, and social, political, cultural, linguistic or modality-specific assumptions. This work audits LLMs as instruments for large-scale framing and perspective analysis in multimodal news coverage. Using an event-centered dataset of 2025–2026 news coverage, where each event includes left-, center- and right-oriented articles about the same headline, we combine embedding-based measures of within-event viewpoint similarity with model-based assessments of framing constructs across modality-specific and metadata-visible input conditions. Rather than treating either dataset labels or model outputs as ground truth, our goal is to examine the usefulness and limitations of LLM-based media analysis. The study contributes an audit protocol that highlights the need to report modality effects, metadata sensitivity, and prompt-induced artifacts alongside substantive claims about news framing and ideological viewpoint differences.

Publication
_The 5th International Workshop on Multimodal Human Understanding for the Web and Social Media (MUWS ‘26), November 10–14, 2026, Rio de Janeiro, Brazil. ACM, New York, NY, USA, 9 pages, https://doi.org/10.1145/3841457.3841515_