Large language models (LLMs) are increasingly used to support Web and social media analysis through model-generated outputs such as labels, scores, summaries, and natural-language explanations or rationales. In many cases, these outputs are used to support broader conclusions about online content, communities, platforms, harm, credibility and social meaning. However, using model-generated outputs as evidence introduces a methodological challenge. These systems can support analysis at scale, but they are not neutral observers, and their judgments may reflect Western-centric norms, culturally specific assumptions, translation artifacts, platform biases and learned stereotypes. In this position paper, we argue that conclusions drawn from LLM-mediated analysis should not be treated as neutral or universal readings of online content, but as outcomes of an analytical pipeline whose outputs depend on language, culture, platform context, modality and interpretation. To support this position, we examine how model outputs are shaped across the analysis pipeline, including task and category definition, data selection, input representation, prompting, model judgment, evaluation, interpretation and use. We discuss these issues through a synthetic multimodal hate speech example. We conclude by outlining implications for more context-sensitive, transparent, and accountable uses of LLMs in Web and social media analysis.