The problem this technology exists to address
Child protection investigation has a volume problem that is difficult to convey. A single seized device can hold hundreds of thousands of files. A single investigation can involve dozens of devices. There is no realistic prospect of a human examiner reviewing all of it, and the material that most urgently needs finding, evidence identifying a child currently at risk, may sit anywhere in that volume.
Automated triage exists to address that directly. Its purpose is to determine what a human looks at first, not to reach conclusions. That framing matters for everything that follows, because the questions raised by a prioritisation tool are quite different from those raised by a system that decides anything on its own.
Hash matching and classification are not the same thing
These two techniques are routinely discussed as one, and conflating them systematically overstates how reliable the newer one is.
Hash matching generates a cryptographic fingerprint of a file and compares it against a database of fingerprints of known illegal material, maintained by bodies such as the Internet Watch Foundation in the UK and the National Center for Missing and Exploited Children in the US. Perceptual hashing variants tolerate minor alterations such as resizing. Critically, this technique does not analyse image content at all: it asks whether this file matches a file already identified by human reviewers. It is long-established, highly reliable, and produces very few false positives.
Classification uses machine learning to assess content in order to flag material that is not already in any database. This is what allows previously unseen material to surface, which hash matching by definition cannot do. It is also where the error rate lives. A classifier produces a probability, not a determination, and it can be wrong in both directions.
The practical consequence is that classifier output is treated as a prioritisation signal for human review, not as a finding. A false positive here is extraordinarily damaging to the person affected, and the safeguard against it is that a human examiner, not the classifier, reaches every conclusion that matters.
The wellbeing dimension
There is a use case here that is rarely discussed outside the profession and is worth stating plainly: reducing how much of this material human beings have to look at.
Investigators and examiners working in this area carry a well-documented occupational burden, and forces have long-standing welfare provisions in recognition of it. Automated triage reduces exposure in two ways. It reduces the total volume a person must review, and hash matching in particular allows already-identified material to be categorised without a human viewing it again. That is a genuine benefit and one of the strongest arguments for the technology, quite separate from any efficiency case.
Generative AI has changed the picture
Generative image models have produced a category of synthetic material that did not previously exist at scale, and it complicates this area in several distinct ways.
Detection is harder, because hash matching relies on a file having been seen and catalogued before, and generated material is novel by construction. Triage is harder, because volume can be produced far faster than it can be reviewed. And investigative prioritisation is harder in a specific and consequential way: a central purpose of this work is identifying children who are currently at risk, and material depicting no real child does not carry that signal while still consuming review capacity.
The legal position varies by jurisdiction and has been actively developing. The UK has legislated in this area, and enforcement questions around generated material continue to be worked through in courts and legislatures internationally. This is one of the faster-moving areas covered on this site, and see deepfakes for the wider synthetic media picture.
Scope creep, the standing objection
The most persistent concern raised about detection infrastructure in this area is not about its use for its stated purpose, which attracts very little objection. It is about extension.
Systems built to detect one category of material can, technically, be pointed at others. Proposals to require client-side scanning on messaging platforms have been the sharpest expression of this argument: supporters view it as the only workable approach to material shared in end-to-end encrypted environments, while opponents argue it creates a general-purpose scanning capability whose future scope depends on policy rather than on technical limits. Both positions are held in good faith by people who agree entirely about the underlying harm, which is worth stating, because this is frequently reported as though one side were indifferent to it.
Follow the coverage
PoliceAI News tracks child protection technology as it develops: detection capability, legislative change, court rulings, platform policy and oversight findings. The feed refreshes every 30 minutes.
View Child Protection Stories