Audio deepfake detection benchmark: deetech.ai takes the top spot

In August, six systems cleared 95% accuracy on this benchmark and an open-weights model reached production grade for the first time. Five new systems later, the board has grown to 23 entries: 13 commercial entries from 12 vendors, and 11 systems whose weights you can download (Pella Research counts in both groups). Eight systems now clear 95%.
Four results stand out in this update:
A new leader. deetech.ai takes first place at 99.56%, with both error rates at 0.4%.
A remarkable debut. Fennura enters in third at 98.63% with a detector it describes as running on CPU alone.
A new speed benchmark. DetectifAI reports the fastest latency of any commercial system, about 24 ms per file.
The open-weights leader holds. Pella Research's pellav2 is still the best model you can download, and still the only one above 95%.
The ground rules have not changed. Vendors routinely self-report around 99% accuracy on their own test sets, and those numbers are not comparable to each other. We hold a fixed set of 4,524 clips with private gold-standard labels and compute every row's metrics ourselves, the same way, on the same files.
The leaderboard

# | System | N | Rej% | Acc% | F1 | FPR% | FNR% | Lat (ms) | RTF |
|---|---|---|---|---|---|---|---|---|---|
1 | deetech.ai (new) | 4524 | 0.0% | 99.56% | 0.996 | 0.4% | 0.4% | 50 | 0.019 |
2 | Resemble DETECT-World | 4524 | 0.0% | 99.47% | 0.995 | 0.7% | 0.4% | 399 | 0.12 |
3 | Fennura (new) | 4524 | 0.0% | 98.63% | 0.986 | 1.2% | 1.5% | 553 | 0.14 |
4 | Aurigin AI § (updated) | 4524 | 0.0% | 98.21% | 0.982 | 2.4% | 1.1% | n/a | n/a |
5 | Resemble AI (DETECT-3B Omni) ‡ | 4524 | 0.0% | 98.05% | 0.981 | 2.5% | 1.4% | 1,164 | 0.40 |
6 | Whispeak | 4524 | 0.0% | 97.70% | 0.977 | 2.9% | 1.7% | 1,052 | 0.39 |
7 | Pella Research (pellav2) † | 4524 | 0.0% | 95.82% | 0.959 | 5.5% | 2.8% | 57 | 0.021 |
8 | Pindrop | 4524 | 0.0% | 95.05% | 0.951 | 6.2% | 3.7% | 282 | 0.076 |
9 | DetectifAI (new) | 4524 | 0.0% | 94.47% | 0.946 | 8.4% | 2.7% | 24 | 0.0084 |
10 | NII Synthetiq Audio v0.8-Beta (new) | 4524 | 0.0% | 89.57% | 0.892 | 6.6% | 14.2% | 91 | 0.024 |
11 | Corsound AI | 3875 | 14.3% | 87.79% | 0.865 | 1.0% | 23.1% | 180 | 0.035 |
12 | Hive ‡ | 4524 | 0.0% | 83.53% | 0.808 | 2.4% | 30.5% | 881 | 0.34 |
13 | Reality Defender ‡ | 3745 | 17.2% | 71.27% | 0.770 | 53.7% | 3.6% | 5,718 | 1.52 |
14 | NII AntiDeepfake † (new) | 4524 | 0.0% | 70.47% | 0.584 | 0.4% | 58.6% | 91 | 0.024 |
15 | Wav2Vec2 (2019 LA) | 4524 | 0.0% | 62.89% | 0.514 | 13.4% | 60.8% | 622 | 0.14 |
16 | AST (ASVspoof 5) | 4524 | 0.0% | 56.83% | 0.657 | 69.0% | 17.4% | 5 | 0.0017 |
17 | Wav2Vec2 (2024 mix) | 4524 | 0.0% | 55.55% | 0.499 | 33.1% | 55.8% | 219 | 0.056 |
18 | Deepfake-V2 (W2V2-base) | 4524 | 0.0% | 53.03% | 0.162 | 3.1% | 90.9% | 94 | 0.027 |
19 | AST (VoxCelebSpoof) | 4524 | 0.0% | 50.99% | 0.048 | 0.5% | 97.5% | 8 | 0.0030 |
20 | RawNet2 (2019 LA) | 4524 | 0.0% | 50.66% | 0.430 | 35.9% | 62.7% | 94 | 0.035 |
21 | LCNN-LFCC (2019 LA) | 4524 | 0.0% | 50.00% | 0.667 | 100.0% | 0.0% | 23 | 0.0056 |
22 | AASIST (2019 LA) | 4524 | 0.0% | 48.17% | 0.486 | 52.6% | 51.1% | 322 | 0.11 |
23 | AASIST3 (ASVspoof 5) | 4524 | 0.0% | 47.63% | 0.029 | 6.3% | 98.4% | 363 | 0.13 |
Bold entries are commercial. N is the number of files actually scored; Rej% is the share a system declined or errored on, so Corsound and Reality Defender are scored on a subset and their rows are not directly comparable to full-coverage rows. † weights are downloadable. ‡ run by Podonos against the vendor's API, with latency measured by us. § replaced an earlier result at the vendor's request (see below). RTF is the mean of per-file latency divided by duration, so short clips weigh heavily.
What changed since August
Five new systems: deetech.ai (new #1), Fennura (new #3), DetectifAI (fastest commercial latency), and two separate systems from the Yamagishi Lab at NII: Synthetiq Audio (licensed commercially) and AntiDeepfake (open weights).
One refreshed entry: Aurigin AI submitted results for a newer model, which replace the run we made against its API in April.
Eight systems now clear 95%, up from six in August and four in June.
A tighter verification statement and a short list of corrections to our August post, both covered below.
A new leader: deetech.ai
deetech.ai takes first place at 99.56%, with a false positive rate and a false negative rate of 0.4% each. That is roughly 1 error in 250 in both directions, with no lean toward over-flagging or under-flagging. It is also the most consistent system across file containers, holding between 99.5% and 99.7% on all six formats, so there is no format it quietly falls over on. And it reports that result at about 50 ms per file.

The lead over Resemble DETECT-World, the August leader, is slim. DETECT-World sits at 99.47%, and on 4,524 clips the 0.09-point gap works out to about four files. It misses fakes at the same 0.4% rate and false-flags slightly more real audio (0.7%). Treat the two as the joint front of the field, and choose between them on integration, latency and price rather than on the accuracy column alone.
The standout debut: Fennura
Fennura is the most striking new entry on the board. It lands in third place at 98.63% on its first submission, ahead of every system from the August update except DETECT-World, with 1.2% FPR and 1.5% FNR. That is a balanced profile: both error rates under 1.6%, in the same company as the two leaders.

What makes the result stand out is how Fennura says it gets there. It describes its system as an on-device detector that runs on CPU alone, with no GPU at inference. If that holds for your deployment, it opens up uses that a hosted API cannot address directly: detection on a handset, at a branch office, or anywhere audio cannot leave the device.
One note for precision. The hardware claim is the vendor's description. We scored the labels Fennura submitted but did not run the system, so we cannot attest to the hardware or to its latency (553 ms per file, self-reported).
The fastest commercial detector: DetectifAI
DetectifAI reports the lowest latency of any commercial system on the board: about 24 ms per file, with a real-time factor of 0.0084. That is roughly twice as fast as deetech.ai (50 ms), about 12 times faster than Pindrop (282 ms), and about 17 times faster than DETECT-World (399 ms). For scoring live calls, streaming moderation or high-volume queues, that headroom matters.

The speed comes with a trade-off worth knowing. At 94.47%, DetectifAI sits just under the 95% bar. It catches fakes well (2.7% FNR) but has the highest false positive rate among systems above 90%, at 8.4%, so about 1 real clip in 12 would be flagged. That fits a fast first-pass filter in front of a stricter check better than a standalone gate.
As with every commercial latency figure outside the three ‡ rows, this one is self-reported on the vendor's own hardware (more on that below).
Eight systems now clear 95%

Alongside deetech.ai, DETECT-World and Fennura, the 95% group looks like this:
Aurigin AI (98.21%) now leans toward catching fakes, at 2.4% FPR and 1.1% FNR. This row is not directly comparable to the one we published before. The earlier result (96.75%, 1.5% FPR, 5.0% FNR) came from an older model that we ran against Aurigin's API in April 2026. The new row is Aurigin's own September 2026 submission for a newer model, and it carries no timing data.
Resemble AI, DETECT-3B Omni (98.05%), the previous generation of DETECT-World, is unchanged.
Whispeak (97.70%) is still the most evenly balanced of the earlier entries, at 2.9% FPR against 1.7% FNR.
Pella Research, pellav2 (95.82%) remains the only open-weights model above the bar.
Pindrop (95.05%) holds its place just over the line.
Below the bar, NII Synthetiq Audio v0.8-Beta comes in at 89.57%, with 6.6% FPR and 14.2% FNR. It is licensed commercially by the Yamagishi Lab at NII rather than sold as a public API.
Pella Research still leads the open-weights board

# | Open-weights system | Acc% | F1 | Licence |
|---|---|---|---|---|
1 | Pella Research: pellav2 | 95.82% | 0.959 | MIT |
2 | NII AntiDeepfake (new) | 70.47% | 0.584 | CC BY-NC-SA-4.0 (non-commercial) |
3 | Wav2Vec2 (2019 LA) | 62.89% | 0.514 | Apache-2.0 |
4 | AST (ASVspoof 5) | 56.83% | 0.657 | BSD-3-Clause |
5 | Wav2Vec2 (2024 mix) | 55.55% | 0.499 | Apache-2.0 |
6 | Deepfake-V2 (W2V2-base) | 53.03% | 0.162 | Apache-2.0 |
7 | AST (VoxCelebSpoof) | 50.99% | 0.048 | MIT |
8 | RawNet2 (2019 LA) | 50.66% | 0.430 | MIT |
9 | LCNN-LFCC (2019 LA) | 50.00% | 0.667 | MIT |
10 | AASIST (2019 LA) | 48.17% | 0.486 | MIT |
11 | AASIST3 (ASVspoof 5) | 47.63% | 0.029 | CC BY-NC-ND-4.0 (non-commercial) |
Licences as stated on each model card. Check the card yourself before relying on this column.
Pella Research's pellav2 is still the best model on this board that you can download and run yourself, at 95.82% under an MIT licence. It is still the only downloadable system to clear 95%, and nothing above it on the main board publishes weights. It leads the next open model by 25.4 points. That gap narrowed from 32.9 points only because a new model moved into second place, not because pellav2 slipped. If you need a model you can inspect, audit or run in your own environment, the practical choice set is still one deep, and pellav2 is it. The weights and an inference script are on Hugging Face.
The new second place is NII AntiDeepfake, an XLS-R 2B model from the Yamagishi Lab (weights on Hugging Face, paper, ASRU 2025), submitted by its authors. At 70.47% it clears the near-random band that every legacy checkpoint sits in. Its error profile is very lopsided: a 0.4% false positive rate, tied for the lowest on the board, and a 58.6% false negative rate. It almost never flags real audio, and it misses close to 6 fakes in 10. That looks like a deliberate operating point rather than a failure to train. Note that its licence is non-commercial.
The August argument holds, and AntiDeepfake adds a second data point. Openness was never the problem; stale training data was. The nine legacy baselines still land between 47.6% and 62.9%. The two open models trained on current synthesis land well above that band.
The real choice is still an error trade-off
With eight systems above 95%, headline accuracy stops separating them. What separates them is which mistake you are buying. A false positive blocks a real customer or takes down legitimate content. A false negative lets a cloned voice through your KYC flow.
Both errors matter roughly equally: deetech.ai (0.4% / 0.4%) and DETECT-World (0.7% / 0.4%), then Fennura (1.2% / 1.5%).
Missed fakes are the expensive error: Aurigin AI's new model (1.1% FNR) and Resemble AI (1.4% FNR) sit just behind the top two.
False alarms are the expensive error: below the 95% group, Corsound AI (1.0% FPR) and Hive (2.4% FPR) rarely false-flag, but miss 23.1% and 30.5% of the fakes they score.
You need to run it yourself: pellav2, accepting a 5.5% false positive rate.
Reality Defender remains the outlier in the other direction. It false-flags 53.7% of the real audio it accepted, which is 1,009 of 1,878 clips. It also declined a further 384 real clips, mostly very short ones, so the false flags amount to 44.6% of all real audio in the set.
Speed, with the same caveat

The latency columns are not verified the way the accuracy columns are. They mix three regimes: the three ‡ rows were run by us against the vendor's API and include network round-trip, other commercial rows are the vendor's own figures on the vendor's own hardware, and the open-source baselines were run locally by us. Treat them as indicative and do not read small gaps as meaningful.
With that said, the new entries push the fast end of the chart further left. Among commercial systems, DetectifAI reports the lowest latency (24 ms), followed by deetech.ai (50 ms) and Pella (57 ms), so the most accurate system on the board is also one of the fastest. NII Synthetiq Audio reports 91 ms, measured by the lab on an H100. Of the 22 systems with timing data, Reality Defender is the only one with an RTF above 1.0 (1.52). Aurigin AI submitted no timing data and is not on this chart.
One reading note. Pella's latency does not scale with clip length, which is the signature of a fixed-length analysis window rather than a full-file read, so its RTF is not comparable to systems that read the whole clip. The AASIST and RawNet2 runners in our repository also score a fixed window of about 4 seconds, taken from the start of the clip and zero-padded when the clip is shorter.
What we verify, and what we do not
This benchmark exists because self-reported numbers are unreliable, so we want to be exact about our own claim.
What we guarantee: we hold the gold labels privately and compute Accuracy, F1, FPR and FNR ourselves for every row. No vendor scores its own row, and the arithmetic behind every number is ours.
What we cannot check: how a submitted predictions.csv was produced. Outside the three ‡ rows (Resemble AI DETECT-3B Omni, Hive and Reality Defender), we did not run the system, so the labels in the file are taken on trust.
The code we once used to call vendor APIs has been retired. Every commercial result added since then, including all five new entries and Aurigin's refresh, comes from a predictions.csv the vendor produced and submitted. That is now the only route onto the board.
Corrections to our August post
While auditing the benchmark repository for this update, with every number recomputed from the gold labels, we found a few statements that were wrong or overstated. Some of them also appeared in our August post:
We wrote that Reality Defender "flags more than half of all genuine audio as fake." The correct figure is 53.7% of the real clips it accepted, which is 44.6% of all real audio in the set, since it declined 384 real clips.
We wrote that accuracy, F1, FPR and FNR were "independently verified for every system." We compute them for every system, but for vendor-submitted rows we did not run the system itself. The section above states this more precisely.
We described pellav2 as scoring a 4-second centre crop, and said our AASIST and RawNet2 runners "crop similarly." We did not run pellav2 and should not have described its internals. Our own runners take a window from the start of the clip, not the centre.
We listed AASIST3's licence as CC BY-NC-4.0. The model card's metadata tag says that, but its licence text says CC BY-NC-ND-4.0, and we now lead with the stricter reading.
We described all ~25 synthesis systems as commercial. Some, including F5-TTS and Chatterbox, are open-weights models we ran locally. Chatterbox is released by Resemble AI, which also appears on this leaderboard.
None of these change a ranking, but they are the kind of detail this benchmark asks vendors to get right.
Telephony tracks: still open
The narrowband (8 kHz, 2G/3G, landline, VoIP) and wideband (16 kHz, 4G/5G) tracks we introduced in August are live. Each contains the same 4,524 clips and hidden labels as the studio set, passed through a randomly assigned speech codec (G.711, GSM-FR, AMR-NB and G.729 for narrowband; EVS-WB and AMR-WB for wideband). No telephony results are published yet. If your detector is meant for call-centre, fraud or KYC audio, this is where it will be tested against the conditions it actually sees, and we would like to see the studio ranking challenged there.
Methodology
4,524 audio clips, balanced 50/50 real and synthetic. Real audio from VCTK (110 English speakers), LJ Speech (single speaker, about 24 hours) and LibriTTS train-clean-360 (about 191 hours, 904 speakers). Synthetic audio from about 25 TTS and voice-cloning systems, a mix of commercial APIs such as ElevenLabs and open-weights models such as F5-TTS and Chatterbox. Every synthetic clip is round-trip transcribed with Whisper before format conversion. Six file formats: .mp3, .wav, .flac, .ogg, .m4a and .webm.
We report accuracy, F1, FPR, FNR, per-format accuracy, rejection ratio, latency and real-time factor. Every system is scored at its default threshold. We deliberately do not report EER, because it assumes an oracle threshold chosen with knowledge of the test labels, which is not a threshold you can set in production.
Submit your system
Run your detector over the public dataset and produce a predictions.csv with four columns: filename, label (exactly real or fake, lowercase), latency_ms and audio_duration_sec. Write NOT_APPLICABLE for files your system declines and error for files it fails on; both are counted in the rejection column rather than against your accuracy. Leave the timing columns blank if you cannot measure them, and they will be reported as n/a.
Submit one CSV per track (studio, narrowband, wideband). The three tracks are shuffled independently, so a prediction file for one will not score on another. Email your CSV to hello@podonos.com. Full details are in the submission format section of the repository.
Limitations
The benchmark is English-only, and it is a snapshot of a fast-moving synthesis landscape. The latency and RTF columns mix measurement regimes. Rows with a non-zero rejection rate are scored only on the files they accepted. Vendor-submitted rows rely on the vendor running its system honestly over the set. And a single accuracy number on a fixed set is the start of a procurement decision, not the end of one.
Run it yourself
The dataset, the open-source runners, the metric definitions and the scoring script are public in the benchmark repository. The gold-standard labels are the one thing we keep back, because a public answer key is how a benchmark stops meaning anything.
Other readings

Audio deepfake detection benchmark: an open model finally clears the production bar
Quickly uncover deep insights into your voice AI's strengths and drive faster development, smarter marketing, and flawless delivery.
|
10 min read

Audio Deepfake Detection Benchmark: the production bar moves to four
Quickly uncover deep insights into your voice AI's strengths and drive faster development, smarter marketing, and flawless delivery.
|
7 min read

Natural ≠ Preferred: What Our TTS Rankings Revealed About How Humans Actually Judge AI Voices
Quickly uncover deep insights into your voice AI's strengths and drive faster development, smarter marketing, and flawless delivery.
|
7 min read

Automatic Deepfake Audio Detection Benchmark
Quickly uncover deep insights into your voice AI's strengths and drive faster development, smarter marketing, and flawless delivery.
|
7 min read

Announcing the Selected Teams for the Podonos Research Support Program
Quickly uncover deep insights into your voice AI's strengths and drive faster development, smarter marketing, and flawless delivery.
|
2 min read

Benchmarking Chatterbox Turbo: How Resemble AI Evaluated Open-Source Voice AI with Podonos
Quickly uncover deep insights into your voice AI's strengths and drive faster development, smarter marketing, and flawless delivery.
|
5 min read

Introducing Podonos Flash
Quickly uncover deep insights into your voice AI's strengths and drive faster development, smarter marketing, and flawless delivery.
|
4 min read

Podonos just raised $2.4M in pre-seed funding
Quickly uncover deep insights into your voice AI's strengths and drive faster development, smarter marketing, and flawless delivery.
|
3 min read

Product Update: Podonos Wizard launch
Quickly uncover deep insights into your voice AI's strengths and drive faster development, smarter marketing, and flawless delivery.
|
2 min read

Why Post-Refining Matters in Voice AI: Making Sense of Raw Evaluation Data
Quickly uncover deep insights into your voice AI's strengths and drive faster development, smarter marketing, and flawless delivery.
|
2 min read