Best AI Music Detector (2026): Eight Tools on One Test Set
We ran eight public AI-music detectors on the same 2,104 audio files at their fixed published thresholds. The honest answer is not a single winner — it depends on whether a false positive (real music flagged as AI) or a missed AI track costs you more. Here is the data, and who each tool actually suits.
Short answer: on this test set, ArtifactNet had the highest F1 (0.952) while FST had the lowest false-positive rate (1.8%). Those are opposite operating objectives. Free to try either trading-off in practice: run ArtifactNet in your browser.
ArtifactBench v2 (new, 2026-09-23): a separate lineage-aware frozen protocol with 828 tracks now runs ArtifactNet at AUROC 0.982 — methodology in the field report, dataset in GitHub · Hugging Face.
How we tested
Every detector received the same restored ArtifactBench v1.1 files: 1,388 AI-generated tracks and 716 real tracks, identical inputs, each scored at its public fixed threshold of 0.5. This is a dated, reproducible snapshot — full methodology, preprocessing adapters and per-source results are in the companion field report and the public result record.
Disclosure: this benchmark was designed and run by Intrect, the maker of ArtifactNet — number one on F1 below. Read the report for the limits before you treat any figure as a purchase decision by itself.
The comparison table
| Detector | Parameters | F1 | Precision | Recall | Real FPR |
|---|---|---|---|---|---|
| ArtifactNet v9.4 ONNX | 4.2M | 0.952 | 0.932 | 97.3% | 13.8% |
| AI-Music-Detection AST-60s | 90.8M | 0.840 | 0.848 | 83.1% | 28.9% |
| CLAM (MoM) | 194.3M | 0.787 | 0.711 | 88.3% | 69.7% |
| SpecTTTra α-120s | 18.7M | 0.777 | 0.880 | 69.5% | 18.4% |
| Deezer ISMIR fakeprint LR | 3.6K | 0.754 | 0.906 | 64.6% | 13.0% |
| FST (Mippia) | 174.4M | 0.735 | 0.984 | 58.7% | 1.8% |
| DeepFense EAT+Nes2Net | — | 0.650 | 0.589 | 72.4% | 97.8% |
| SpecTTTra β-5s | 18.7M | 0.563 | 0.884 | 41.3% | 10.5% |
Why you should care about false positives
A label, distributor or moderator reviewing a catalog does not only pay for missed AI tracks. Every genuine release incorrectly flagged consumes review time and can damage a creator's standing. In this run two systems with seemingly useful recall flagged more than two-thirds of real music — recall alone is a trap. FST made the opposite trade: the fewest false alarms, but it missed far more AI tracks.
The right tool is the one whose error profile matches your volume and tolerance. If you review thousands of tracks and one false positive is a real fight, a low-FPR tool wins. If you need to catch high volumes of AI music and can spend human time on triage, recall-and-precision balance matters more.
Who each tool suits
ArtifactNet
Highest F1 with 97.3% recall at a manageable 13.8% FPR, and it is the only lightweight model here (4.2M params) with a free in-browser demo. Good default if you want the strongest overall balance. Try it free: demo.intrect.io.
FST (Mippia)
1.8% false-positive rate on real music — the safest choice when a wrong flag is expensive. You trade recall (58.7%): it will let more AI tracks through than the leaders.
AI-Music-Detection AST-60s
Solid precision (0.848) and recall (83.1%). A reasonable middle option if ArtifactNet's FPR is more than your operation tolerates.
CLAM / SpecTTTra
Useful research references. CLAM catches AI sections but flags ~70% of real music in this run; SpecTTTra trades recall for a lower false-alarm rate. Both are heavier.
What this comparison does not prove
- The restored real subset was 716 available files, not every row in the purged manifest — see the report's restoration boundary.
- A fixed threshold is reproducible but not necessarily optimal for your error budget.
- Generators and codecs change; a detector tuned to today's Suno/Udio may lag a new model. Query the maker for the generation they cover before you commit to volume.
- Detection is not provenance. No score replaces Content Credentials, platform metadata, or a human review-and-appeal path.
What detector can't do alone
Detecting AI music answers where a track may have come from. Fixing the audible artifacts it leaves on the audio is a separate job — this is what the de-artifact plug-in and online cleanup handle by removing RVQ ghosting and codec residue.
Sources and revision policy
Primary sources: the 8-way field report, the ArtifactBench dataset card, the public result table, and the ArtifactNet paper. This is a dated snapshot: a rerun gets a new report or an explicit correction, never a silent overwrite.
Try the top-F1 detector on your own file
Upload a track and get an AI-vs-human verdict with forensic evidence — free, in the browser, no account, nothing retained.
Run a free AI-music detection analysis
First analysis is free. Treat the result as forensic evidence to review, not an automatic verdict.
Run a free analysis →