Cognitive biases are commonly defined as systematic deviations from rationality in human judgment and decision making. They have been studied in psychology for decades, with corresponding literature claiming to have found evidence for several hundreds of such biases. Driven by the success of large language models (LLMs), researchers have recently begun to investigate whether cognitive biases also emerge in LLMs. While some evidence has been found, respective studies are typically limited to textual interactions with the LLM under investigation, i.e., providing a text prompt aimed at uncovering a cognitive bias and receiving a textual answer. Despite today’s widespread adoption of vision–language models (VLMs) and multimodal large language models (MLLMs), multimodal investigations of cognitive biases in generative artificial intelligence systems are rare. Addressing this gap, in the position paper at hand, we review cognitive biases identified in psychology whose study in LLMs requires multimodal data. Focusing on the text and image modalities, we discuss the identified biases, offer corresponding research hypotheses, and provide preliminary evidence.