Independent · Updated continuously
Artificial Intelligence in Law Enforcement
Child protection

CSAM Detection and Classification

Deep-learning classifiers that detect previously unseen child sexual abuse material, including deepfake CSAM, and grade its severity — going beyond hash-matching of already-known images.

Various; UK deployment built by Roke

National child abuse image databases have historically relied on hash matching: comparing a file's digital fingerprint against fingerprints of already-identified abuse material. That catches what is known and nothing else.

Classification is the newer capability. Deep learning models trained on an existing corpus detect characteristics of abuse imagery in files never previously catalogued, and grade severity against standard government categories — work that would otherwise require an officer to view the material directly.

THE UK IMPLEMENTATION

The Child Abuse Image Database went live in December 2014 and is used by all UK territorial forces, Police Scotland, the PSNI and the National Crime Agency. It holds images, videos, case reference data, data about faces and objects appearing in material, grading data, hash values and victim characteristics — but not victim names or contact details.

Content arrives from suspect devices seized in raids, from overseas forces and agencies, and from the Internet Watch Foundation, which has had access since 2018 and can upload directly.

The AI layer is the Vigil AI CAID Classifier, developed with the Home Office CAID team and delivered by Roke. Roke describes it as, to its author's knowledge, the first operationally viable CSAM classifier able to both detect and grade, and states it handles previously unknown material including deepfake CSAM. Those claims are attributed to the supplier rather than stated as established.

The Home Office had invested £18.2 million by 2019, with AI and fast-forensic innovations costing a further £1.76 million, developed with UK suppliers including Qumodo, Vigil AI and Cyan Forensics. The database held 13 million images at that point, growing by around half a million every two months.

THE CASE FOR IT

The wellbeing argument is the strongest and is rarely made explicit elsewhere in this catalogue: automated grading reduces how much abuse material a human officer must personally view in order to categorise a case. That is a genuine benefit to the people doing this work, whose exposure has documented psychological consequences.

The capability gap is also real. Hash matching cannot identify a newly produced image, which means it cannot help identify a child currently being abused. A classifier can flag material no one has seen before, and that is the difference between historic and live safeguarding.

The deepfake capability, if it performs as described, addresses a category of harm that is growing and that hash matching cannot touch at all.

THE CASE AGAINST

The claims are supplier claims. Roke's white paper is a commercial document, and no independent evaluation of the classifier's accuracy — its false positive or false negative rate, or whether performance varies by the apparent ethnicity or age of the child depicted — has been identified.

The consequences of error are asymmetric and severe in both directions. A false negative leaves material uncatalogued; a false positive attaches the most serious possible allegation to an individual. Grading errors affect sentencing.

There is also no dedicated statute or named external regulator for CAID, which for a database of this sensitivity is a notable gap. A Home Office market notice in May 2024 signalled a possible move to public cloud hosting, reversing a 2020 legal opinion that data protection concerns meant the data should not be cloud hosted; the hosting contract expired in early 2026 and the outcome is not established.

WHAT IS NOT ESTABLISHED

No independent accuracy evaluation of the classifier has been identified.

Whether the cloud hosting move proceeded is not established.

Current database size and growth rate are not published beyond the 2019 figures.

Whether the deepfake detection capability has been tested against current generative models, which have advanced substantially since the tool was built, is not documented.

Related subject: Child Protection and AI

Where this is deployed

Full tracker →
CountryForceStatus
UKAll UK territorial forces, Police Scotland, PSNI and the NCAUK-wideOperational
SGSingapore Police ForceSingapore (national)Operational

Sources

  1. GOV.UK CAID privacy notice; Roke white paper; Computer Weekly — All UK territorial forces, Police Scotland, PSNI and the NCA