Questioning the Effectiveness of Person Re-Identification Datasets
14 May 2026, by Viktoria Wrobel
Mass surveillance technologies rely on person re-identification (ReID), the task of matching the same person across non-overlapping cameras. ReID datasets are often built from footage of pedestrians, students, and shoppers, recorded by surveillance cameras in public space without their consent. These datasets have drawn sustained ethical criticism, yet they remain openly hosted and widely used, defended by the claim that the models trained on them are useful for surveillance deployment. We test that claim on the tracking task the footage was collected to enable, and it does not hold. We find that training on ReID datasets does not improve tracking over a general pretrained encoder, across eight architectures and three pedestrian-tracking benchmarks, and that pooling five datasets into one larger training set does not change this. We further find that this training produces representations that encode a shortcut, clothing colour, rather than identity: clothing colour is far more decodable from these representations than identity, and different people wearing the same colour produce three to five times as many false matches. We add empirical depth to the ethical case against ReID datasets, which until now has rested on the consent violation alone. With no benefit left to set against that violation, we argue these datasets should not be built.
This basecamp project was carried out by Pranav A and the results will be published in AIES 2026.

