Overview and Joint Report of the Voight-Kampff Task of the ELOQUENT 2026 Lab for Evaluating Generative Language Model Quality: Notebook for the ELOQUENT Lab at CLEF 2026

Abstract

The challenge for generative AI authorship verification is not simply distinguishing human-written text from machine-generated text, but doing so under increasingly realistic and adversarial conditions. The 2025 Voight-Kampff task emphasized that obfuscated AI text and human-AI collaborative writing remain difficult cases, with detection performance affected by domain shifts, rewriting strategies, and subtle differences between machine-polished, machine-humanized, and human-edited content. This year, we built directly on this problem by focusing on whether generated texts can become harder to detect across varied sources, genres, and generation strategies. This year’s competition reflects that shift. Across five participating teams and 16 submitted responses, systems differed not only in model choice but also in how they attempted to create human-like text: through prompt engineering, controlled decoding, detector-aware generation, and fine-tuning. The results are explained in the end, which reflects a more diverse experimental setting that helps reveal where current detection systems remain robust and where they begin to fail.

Publication
CLEF 2026 Working Notes, September 21–24, 2026, Jena, Germany