Training seeds and model-selection stability in recommender-system evaluation

Abstract

Recommender-system experiments often report results from a single random training seed, assuming that run-to-run stochasticity has little effect on conclusions. This study tests that assumption by fixing the data split and varying training seeds across hyperparameter configurations. Seed effects are evaluated through user-level metric sensitivity, validation-based model selection, and agreement between recommendation lists. The results show that seed variation is often detectable and that its impact depends on how clearly configurations are separated, how well validation behavior transfers to the test set, and whether similar scores produce similar top-k lists. The findings argue that training seeds should be treated as part of the evaluation protocol rather than incidental implementation noise.

Publication
_Proceedings of the 20th ACM Conference on Recommender Systems (RecSys 2026), Minneapolis, United States, https://doi.org/10.1145/3773078.3841289_