Skip to content

AI

Seeing data leakage: duplicate images in an MRI dataset with pHash

High accuracy isn't always good news. How I found duplicate images and how cleaning them changed the results.

Published:
Reading:
1 min read

Introduction

While working on NeuroVision AI, the first results were too good. When a classifier makes almost no mistakes, it is usually telling you a story about the data rather than the model. In this post I explain how I looked at the data before trusting the result.

The problem: same image, two sets

Public MRI datasets are assembled from several sources. The same image can appear more than once under a different name, slightly cropped or resized. With a random split, these copies end up in both the training and the test set.

The result: at test time the model meets an image it has already seen, and the reported accuracy comes out higher than its real ability to generalise.

What is pHash?

A perceptual hash is a short fingerprint computed from an image's structure rather than its raw pixel values. SHA-256 catches exact copies, but a resized or recompressed copy has a completely different SHA value. pHash, on the other hand, produces very close values for such images.

Deterministic data preparation

To keep the comparison fair, all five architectures used the same cleaned data and the same split:

# Every experiment shares the same split
SEED = 42
df = dedupe(df, exact="sha256", perceptual="phash")
gss = GroupShuffleSplit(test_size=0.2, random_state=SEED)
train_idx, test_idx = next(gss.split(df, groups=df.group_id))
 
# Leakage check: no group may appear in both sets
assert set(df.group_id[train_idx]).isdisjoint(df.group_id[test_idx])
  • Duplicate cleaning: SHA-256 first, then pHash
  • Group-based split: GroupShuffleSplit
  • Leakage check: assertions built into the code

What changed?

Before cleaning, the weakest class was meningioma. After cleaning the picture changed and the weakest class became healthy (no tumour) images. The leak was limited but real; the reported results now rest on the cleaned data.

Conclusion

Question the data before changing the model architecture. Duplicate cleaning is a few lines of code, but it completely changes how much you can trust the results. For the full project, see the NeuroVision AI case study.

  • Computer Vision
  • Data leakage
  • pHash
Share

This post in other languages: Türkçe · العربية