The system placed 25 people at a protest. Prison records placed them in jail.
An investigation checked 205 of the 2,873 people Delhi Police told India's Supreme Court its facial recognition system had identified at the Jantar Mantar protests. Prison records placed at least 25 of them in custody at the moment they were flagged. What makes the finding unusual is not the number. It is that anyone was able to check at all.
Between 20 and 26 July 2026, Delhi Police used facial recognition at the Jantar Mantar protests over a leaked examination paper. In an affidavit filed with India's Supreme Court on 17 August, the force said the system had identified 2,873 people with criminal antecedents at the protest site.
On 4 September, The Indian Express published the result of checking a sample of that list against prison records. Of 205 people it examined, at least 25 were in custody in Delhi's Tihar, Mandoli or Rohini prison complexes at the moment the system placed them at the protest.
The obvious reading is that a facial recognition system made 25 mistakes. That reading is probably too simple, and the more interesting point sits somewhere else entirely.
What the police told the court
The affidavit broke the figure down. Of the 2,873 identifications, 2,402 came from the police biometric database known as Crime Kundli, and 471 came through criminal records. The force stated that no action would follow from a match alone, and that field verification would be carried out first.
On 1 September the Supreme Court quashed the first information reports registered in connection with the student protests, while allowing the authorities to proceed against the 2,873 people the police said had criminal records, subject to that verification. The court is separately considering whether the use of facial recognition and other biometric tools at the protests had a lawful basis at all, and whether it was compatible with rights to privacy, expression and peaceful assembly.
So the list is not an abstract intelligence product. It is a list the state has been permitted to act on.
What was actually checked
The sample was not random, and it matters that it was not. The newspaper selected 205 people from the list who were linked to the most serious offence categories: murder, attempted murder, rape and offences under the Protection of Children from Sexual Offences Act. Those are the entries where an error would carry the heaviest consequence.
Within that group, the 25 found to be in custody comprised 17 people facing murder charges, four facing rape or other sexual offence charges including two POCSO cases, and four facing attempted murder charges. Three were flagged on 24 July, twenty-one on 25 July and one on 26 July.
Twenty-five out of 205 is not a false positive rate. The sample was chosen deliberately rather than drawn at random, the figure is a floor rather than a total, and nothing about the remaining 2,668 entries has been established either way. Anyone quoting this as a percentage error rate for the system is misusing it.
What it does establish is that the list contains entries that cannot be correct as stated, in a category where the force said the antecedents were most serious.
The rare thing here is the check, not the error
This is the part worth dwelling on.
Almost no facial recognition deployment anywhere can be audited the way this one was. The reason this finding exists is that Indian prison records constitute an independent dataset, held by a different institution, recording where a specific person physically was at a specific time. A journalist could hold the police list against it and see where the two disagreed.
Now consider how accuracy is normally established for these systems. A force publishes the number of alerts its system generated and the number it classifies as incorrect. In the United Kingdom, forces publishing live facial recognition deployment figures routinely report incorrect alert counts of zero.
Those figures are not fabricated. But they are produced entirely inside the system that produced the alerts. An alert becomes an incorrect alert when an operator or an officer decides it was wrong, usually at the moment of engagement. If the officer approaches the right person, or approaches nobody, or approaches someone who does not contest it, the alert is not recorded as an error. Nothing external is consulted.
That is not an accusation of bad faith. It is a description of what the measurement can and cannot see. A self-adjudicated error count captures the errors that were noticed. It cannot capture the errors that were not.
Delhi's 25 were caught precisely because someone stepped outside that loop. It took access to prison, police and court records, and a newspaper willing to work through them one name at a time.
What "identified" is doing in that sentence
Two further details from the record are worth setting alongside the finding.
The first concerns thresholds. In a response to a right to information request in 2022, Delhi Police indicated that it treated an 80 per cent match as a positive result. A similarity score of 80 per cent is not a statement that the identification is 80 per cent likely to be right. It is a measure of how closely two facial templates resemble each other under whatever conditions the image was captured in. The relationship between that score and the probability of a correct identification depends on the algorithm, the database, the image quality and the size of the population being searched. Delhi Police has not published a false positive rate at its operating threshold.
The second concerns who gets scanned. Appearing before the Supreme Court, the Solicitor General said the system "does not identify everybody indiscriminately", describing it as checking captured faces against records of people with serious criminal antecedents.
Technologists have questioned whether that description can hold. A live system cannot know in advance which faces belong to people on a database. It has to detect every face in view and generate a biometric template for each one before any comparison is possible. What varies between systems is how quickly non-matching templates are discarded, not whether they are created.
That distinction is not pedantry. It determines whether the operation is best described as checking a watchlist or as biometrically processing everyone present at a protest and retaining the subset that matched.
What the finding does not show
Being precise about this cuts both ways, and there are at least three explanations for the 25 that do not involve the face-matching algorithm failing.
A match may have been correct as a face and wrong as an identity, if the database record attached to that face carries the wrong name or has been linked to the wrong criminal history. That is a data quality failure, not a recognition failure, and it has a different fix.
The list may reflect people named in criminal cases rather than people convicted, and the connection between a name, a record and a photograph in a large policing database is not always sound. Custody records themselves can carry errors.
And the police position remains that these are unverified leads, with field verification pending across all 2,873 entries.
None of those explanations is reassuring. Each one describes a different way for a person to end up on a list the state has been authorised to act upon, when they were demonstrably somewhere else. But they point to different remedies, and a serious account has to say which failure it is alleging. On the public record so far, nobody can.
Field verification is the safeguard, and it is unexamined
The safeguard offered in every one of these systems is the same: a human being checks before anything happens to anyone.
The difficulty is that the safeguard is nearly always described rather than evidenced. Delhi Police says verification of all 2,873 identifications is pending. What is not established is what verification consists of, what proportion of matches it rejects, who performs it, or whether its outcomes are recorded in a form that anyone outside the force could later examine.
A verification step that rejects a substantial share of matches is a functioning control and its rejection rate is the single most informative number about the system. A verification step that rejects almost nothing is a formality. From outside, the two are indistinguishable, and the same is true of the human review step in most police AI systems now in operation.
Why this matters beyond Delhi
Facial recognition at protests is spreading. The Metropolitan Police used it in a protest policing operation in London for the first time in May 2026. India's Supreme Court is now weighing the legality of its use at Jantar Mantar directly.
The argument made for these deployments is that the technology only surfaces people already known to the police, and that human judgement governs what follows. The Delhi finding does not refute that argument. It does show what happens when someone is finally able to test one of its outputs against a record kept somewhere else.
Which suggests the question that ought to be asked of every operator of this technology, in any country: what independent record exists against which your matches could be checked, and has anyone ever checked them against it?