COMPAS and US Pretrial Risk Assessment
Actuarial tools scoring a defendant's likelihood of reoffending, used in US bail, sentencing and parole decisions, and the subject of the most consequential algorithmic fairness dispute yet conducted.
Equivant (formerly Northpointe) for COMPAS; Arnold Ventures for the Public Safety Assessment Supplier's own site ↗
COMPAS — Correctional Offender Management Profiling for Alternative Sanctions — scores a defendant's likelihood of reoffending from factors including criminal history and personal circumstances. It was built by Northpointe, a company founded in the 1990s, which merged with two sister companies to become Equivant in January 2017. The scores inform bail, sentencing and parole decisions across many US jurisdictions.
It matters far beyond its own deployment because the argument about it produced the clearest public demonstration of something that turns out to be mathematically unavoidable in every risk-scoring system on this site.
THE PROPUBLICA FINDING
In May 2016 ProPublica published "Machine Bias", analysing COMPAS scores for more than 10,000 defendants in Broward County, Florida.
It found the tool correctly predicted reoffending about 61% of the time. More importantly, it found the errors fell unevenly: Black defendants were almost twice as likely as white defendants to be labelled higher risk and then not reoffend, while white defendants were considerably more likely to be labelled lower risk and then go on to offend. The mistakes ran in opposite directions by race.
THE COMPANY'S REBUTTAL
Northpointe responded with a 37-page technical document, and its argument was not a denial. The company accepted the overall accuracy figure of around 60%, and demonstrated — using ProPublica's own data — that this accuracy was equivalent for Black and white defendants. A defendant assigned a given risk score had roughly the same actual probability of reoffending regardless of race.
Its criticisms of the analysis were specific: that ProPublica had not accounted for different underlying base rates of reoffending between the groups, and that combining the High and Medium categories into a single "higher risk" band inflated the apparent false positive rate.
WHY BOTH SIDES ARE RIGHT
This is the part that generalises, and it is why this page matters more than its deployment count suggests.
ProPublica measured whether error rates were equal across groups — whether a Black defendant and a white defendant who both did not reoffend were equally likely to have been wrongly labelled high risk. Northpointe measured whether a given score meant the same thing for both groups — whether a high-risk label carried the same actual probability of reoffending.
Both are reasonable definitions of fairness. Both were satisfied by the evidence each side presented. And subsequent academic work established that where two groups have different underlying base rates, a scoring system cannot satisfy both simultaneously. It is not a flaw in COMPAS that better engineering removes. It is a property of the mathematics.
That has a direct consequence for every risk tool in this catalogue: a vendor claiming its system is fair is claiming something incomplete unless it says which definition of fairness it means, and what it therefore gives up.
WHAT THE COURTS DID
In State v. Loomis, decided by the Wisconsin Supreme Court in July 2016, Eric Loomis challenged the use of a COMPAS score in his sentencing as a violation of due process, arguing he could not examine a proprietary algorithm. A judge had cited the assessment in describing him as high risk to the community before imposing eight and a half years.
The court rejected the challenge. Its reasoning included that identical COMPAS reports were available to both the defendant and the state — that is, neither side could inspect the algorithm, so neither was disadvantaged relative to the other. The court permitted use of the score while cautioning against reliance on it as a determining factor.
THE CASE FOR IT
The honest comparison is not between an algorithm and perfect justice; it is between an algorithm and unaided human judgement, which carries its own well-documented and less measurable biases. A structured tool applies the same factors to every defendant, which is more than can be said for intuition, and it can be audited — as it was.
Nothing in the ProPublica analysis showed the tool was worse than the judges it advised. That comparison has largely not been made.
THE CASE AGAINST
Sixty-one percent accuracy is the figure that should give most pause, and it is not disputed by either side. A tool that is right about three times in five is informing decisions about liberty.
Proprietary opacity remains unresolved. Loomis reasoned that equal access to the same limited information satisfied due process; it did not address whether anyone can meaningfully contest a score whose workings nobody outside the vendor can inspect.
The feedback problem applies here as it does to place-based prediction: a model trained on rearrest learns patterns of enforcement as much as patterns of offending.
And there is a use question distinct from the accuracy question. Northpointe's founder later indicated that race-correlated factors may have been in play, and that he never intended COMPAS to be the sole basis for a decision. A Broward County judge who oversaw most pretrial hearings said he had stopped relying on it years earlier, preferring his own judgement. How a score is weighed sits outside the vendor's control and is rarely documented.
WHAT IS NOT ESTABLISHED
How many jurisdictions currently use COMPAS or comparable tools, and in what decisions, is not centrally published.
No study identified compares outcomes under COMPAS-informed decisions against unaided judicial decisions on equivalent cases, which is the comparison that would actually answer whether it helps.
Whether the tool has been revised since 2016 in response to the findings, and whether any revision has been independently evaluated, is not established.
How often judges depart from a score, and in which direction, is not systematically recorded.
Where this is deployed
Full tracker →| Country | Force | Status |
|---|---|---|
| US | US state courts and corrections agenciesMultiple US states | Operational |
| US | US pretrial courts (multiple jurisdictions)Nationwide | Operational |
Sources
- ProPublica — Machine Bias, the original investigation
- ProPublica — how we analyzed the COMPAS recidivism algorithm (methodology)
- Equivant — vendor site (formerly Northpointe)
- Colorado Technology Law Journal — how to argue with an algorithm: the COMPAS-ProPublica debate and State v. Loomis
- UCLA Law Review — predictive algorithms in criminal sentencing, acknowledging both sides' figures
- Columbia Human Rights Law Review — competing fairness definitions in the COMPAS dispute
- arXiv — technical treatment of the Northpointe rebuttal and the fairness definitions in conflict