Independent · Updated continuously

PoliceAI News

Artificial Intelligence in Law Enforcement
Biometrics

Phonexia Voice Biometrics

Czech speech technology used by police and forensic units to identify speakers by voice across recordings, bundled with transcription, language identification, deepfake detection and contested emotion and age estimation.

Phonexia (Brno, Czech Republic) Supplier's own site ↗

Phonexia is a Czech speech technology company, based in Brno and founded in 2006 as a spin-off from Brno University of Technology. Its Speech Platform bundles fourteen modular technologies covering voice biometrics and speech recognition, and it sells primarily to government, law enforcement and forensic units, reporting a presence in more than 60 countries.

Its core capability is speaker identification: determining whether the voice in two recordings belongs to the same person, independent of what was said, in what language, or over what channel.

HOW THE COMPANY DESCRIBES IT

Phonexia frames speaker identification around two distinct questions, and the distinction is a useful one. Identification asks "whose voice is this?", typically as a one-to-many or many-to-many comparison — the company gives the example of fake emergency calls, and states that large-scale automatic identification is used by law enforcement for database searches and ranking suspects. Forensic voice analysis is the narrower later-stage task, using smaller amounts of data and one-to-one comparison to evaluate evidence, establish the probability of a speaker's identity, and produce something usable in court.

The technical premise, in the company's own account, is that speech organs and speaking habits are more or less unique to a person, so the features of a captured speech signal are also more or less unique. The hedging in that phrasing is the vendor's own, and is more careful than most biometric marketing.

Phonexia claims over 98% accuracy in evaluations, participates in the NIST Speaker Recognition Evaluations, and works with a NIST-recognised speech research group at Brno University of Technology. Its platform is built exclusively on deep neural networks, extracting around a hundred feature vectors per second of audio.

Deployment is deliberately on-premises or private cloud, with the company stating that sensitive data never leaves the customer's servers. Phonexia also notes it is a private Czech company with a transparent ownership structure, governed by EU legislation — a positioning aimed at buyers wary of biometric suppliers from other jurisdictions.

WHAT ELSE IS IN THE PLATFORM

Beyond speaker identification, the suite includes speaker diarization (separating and labelling individual speakers in a single recording), language identification across 140 languages, transcription in more than 60, keyword and phoneme spotting, voice activity detection, speech quality measurement, voice deepfake detection, and estimation of a speaker's age, gender and emotional state.

WHERE IT IS USED

Deployments recorded on this site's tracker are listed below. The Czech Police names Phonexia among the tools it classifies as applying machine learning, in its own freedom-of-information disclosure.

Phonexia was also the voice biometrics provider in ROXANNE, an EU Horizon 2020 project bringing together 25 partners across 15 countries, including 11 law enforcement agencies, to build a platform for investigating organised criminal networks. The resulting Autocrime platform combined speaker identification with multilingual speech recognition, gender identification, keyword and topic detection, named entity recognition and network analysis, and the consortium said it would be offered to European law enforcement under an open-source licence. The project also produced ROXSD, a synthetic but realistic dataset representing a fictional organised crime network, released for research use.

THE CASE FOR IT

Speaker identification addresses a genuine investigative problem. Intercepted or recorded audio frequently features unidentified voices, and comparing recordings by ear is slow and unreliable. A system that ranks candidates from a voiceprint database gives investigators a starting point rather than an answer.

The company's separation of investigative identification from forensic one-to-one comparison is methodologically sound, and it matters: using a database search result as a lead is defensible, whereas presenting the same output as courtroom evidence is not, and Phonexia's own documentation draws that line explicitly.

On-premises deployment is a real advantage for this class of data. Voice recordings from investigations are among the most sensitive material a force holds, and a model where audio never leaves the customer's infrastructure avoids an entire category of risk that cloud-hosted biometrics carry.

The synthetic ROXSD dataset deserves credit too: releasing realistic training data that contains no real people's communications is a better answer to the research-data problem than the alternatives.

THE CASE AGAINST

Two capabilities in the wider suite warrant more scrutiny than the speaker identification does.

Emotion recognition from voice is scientifically contested. The claim that a machine can identify a speaker's emotional state from acoustic features is not settled science, and the EU AI Act treats emotion recognition as a restricted category in several contexts precisely because of that. Its inclusion in a law enforcement platform, alongside age and gender estimation, invites inferences about people that the underlying evidence may not support.

The detention use case is the more striking. Phonexia's own material describes detecting suspicious behaviour in detention centres including "misuse of call allowances or use of prohibited languages" to ensure inmates comply with phone rules. Automated detection of which language a detained person is speaking, in order to enforce a prohibition on speaking it, is a significant capability to see described as a routine product feature, and it sits close to the concerns raised elsewhere in this catalogue about continuous monitoring of prisoner communications.

The accuracy claim also needs reading carefully. Over 98% in evaluations does not specify which evaluations, at what dataset size, or under what audio conditions. Speaker identification performance degrades with short samples, background noise, poor channels and deliberate disguise, and none of those are addressed by a single headline figure.

Finally, one-to-many identification depends on a voiceprint database existing. Building one means enrolling voices, and the governance around whose voice is enrolled, on what basis, and for how long, is not addressed by the technology and is rarely addressed publicly by the forces using it.

WHAT IS NOT ESTABLISHED

Which specific evaluations produce the 98% figure, and under what conditions, is not published.

No independent assessment of the emotion recognition or age estimation components was identified, nor any statement of whether police customers enable them.

How voiceprint databases are governed in any deployment reviewed — enrolment criteria, retention, deletion — is undocumented.

Whether the Autocrime platform from ROXANNE was adopted operationally by any of the 11 participating law enforcement agencies after the project concluded is not established by the sources reviewed.

Related subject: Biometrics and Identification

Where this is deployed

Full tracker →

No tracker entries currently reference this technology.

Sources

  1. Phonexia — voice biometrics and speaker identification (vendor)
  2. Phonexia documentation — speaker identification, forensic voice analysis and the 1:1 / 1:N distinction
  3. Phonexia — Speech Platform for Government, including the detention centre use case
  4. Phonexia — Speech Platform modular technology suite
  5. Biometric Update — EU ROXANNE consortium and the Autocrime platform
  6. Biometric Update — Phonexia forensic voice biometrics product launch
  7. Policie CR — freedom-of-information disclosure naming Phonexia among its machine learning tools