Plurall AI Timbre

Audio in, a calibrated likelihood out. Every number on this page was measured before the page existed, and the rates this detector is wrong at are shown beside the verdict rather than behind it.

Drop an audio file here

or

wav, flac, mp3, ogg

What happens to the file

It is sent to our scoring service over an encrypted connection, decoded, scored, and deleted. Nothing is retained: the log records a content hash, the duration and the verdict, and never the audio or the filename. Audio shorter than one second, effectively silent, or below an 8 kHz source rate is refused rather than scored.

What it does, measured

Each figure names the document in the record that carries it. A detector is only as good as the conditions it was measured under, so the conditions are stated rather than summarised.

In the domain it was built for

Worst held-out generator condition
Six generator-disjoint and speaker-disjoint folds. The headline is the worst fold, not the mean of 1.46 %.
3.16 % EERdocs/97
Reserved non-vocoder condition
One attack built without a neural vocoder, held out of training entirely and scored in every fold. One generator, and the same clips each time.
AUC 0.9999docs/51
Held-out equal error rate of the shipped head
On the partition the operating points were derived from.
0.293 %docs/97

Against current commercial synthesis

Five Google Cloud TTS voice families
Full-length audio measurably at the training channel, the newest family missed most completely. Three engines, not five: three of the families return numerically identical audio. The finding is docs/72; this is the figure re-measured under the shipped configuration.
93.45 % to 100.00 % misseddocs/97
Genuine half of that same corpus
The control. The failure is the generator, not the channel.
0.13 % false alarmsdocs/72
A different vendor's generator, held out
At the shipped threshold, after retraining on one vendor. 93.01 % at a matched 1 % false-alarm point.
84.62 % misseddocs/97
Cross-vendor transfer
At matched 1 % false-alarm operating points. Teaching the detector one vendor buys almost nothing against another. One held-out vendor, so one data point.
8.9 % of the within-vendor benefitdocs/98

On audio it was not built for

Read speech, different source family
Full-length clips at the 1 % in-domain operating point.
0.16 % false alarmsdocs/97
Spontaneous speech, consumer microphones
Same threshold, same measurement. The ordering across corpora is unexplained: acoustic distance from the training channel was tested as the cause and rejected.
20.51 % false alarmsdocs/97

On an independent public benchmark

ASVspoof 2019 LA evaluation set
All 71,237 trials under the shipped configuration. Computed here under a published definition and positioned against a published column; nobody else scored it. The comparable system on that column, the only other one not trained on that corpus, is at 5.833 %.
19.391 % EERdocs/98

Short audio, and what changed

Genuine one-second clips, zero-padded
The previous preprocessing. Speech content held fixed, the fill the only variable, so this is causal rather than correlational.
92.40 % falsely flaggeddocs/68
The same clips, tiled to fill the window
What ships. Audio under one second is refused rather than scored.
2.67 % falsely flaggeddocs/68

Operational

Single-request scoring latency
Batch 1, warm, server-side: read, resample, encoder and head, excluding network. Measured on the deployment hardware over 20 requests, 27.7 to 28.2 ms.
27.9 ms mediandocs/71
Audio it refuses rather than guessing on
A detector that returns a number for input it cannot judge is worse than one that declines, because the number will be used.
under 1.0 s, silent, below 8 kHzdocs/99

3 further measured figures are withheld from this page. They were measured on a corpus whose licence is recorded as conflicted, and whether a figure measured on it may be published is an open question with counsel. They are in the internal record and are not withheld from it.

What it does not do

It does not detect current commercial synthesis, and that is the most serious limitation here rather than a footnote. It does not generalise across vendors: teaching it one vendor's generator buys almost nothing against another. It scores one four-second window, so nothing beyond 22.4 seconds has been measured in any configuration. And it was trained on the development partition of its source corpus, which places it outside the condition published challenge results use, so no figure here is comparable to a published challenge result.