Intrect / Artifact Observatory / Comparison guide
Buyer's comparison · Intrect Research

Best AI Music Detector (2026): Eight Tools on One Test Set

We ran eight public AI-music detectors on the same 2,104 audio files at their fixed published thresholds. The honest answer is not a single winner — it depends on whether a false positive (real music flagged as AI) or a missed AI track costs you more. Here is the data, and who each tool actually suits.

Published September 23, 2026Intrect ResearchArtifactBench v1.1

Short answer: on this test set, ArtifactNet had the highest F1 (0.952) while FST had the lowest false-positive rate (1.8%). Those are opposite operating objectives. Free to try either trading-off in practice: run ArtifactNet in your browser.

ArtifactBench v2 (new, 2026-09-23): a separate lineage-aware frozen protocol with 828 tracks now runs ArtifactNet at AUROC 0.982 — methodology in the field report, dataset in GitHub · Hugging Face.

How we tested

Every detector received the same restored ArtifactBench v1.1 files: 1,388 AI-generated tracks and 716 real tracks, identical inputs, each scored at its public fixed threshold of 0.5. This is a dated, reproducible snapshot — full methodology, preprocessing adapters and per-source results are in the companion field report and the public result record.

Disclosure: this benchmark was designed and run by Intrect, the maker of ArtifactNet — number one on F1 below. Read the report for the limits before you treat any figure as a purchase decision by itself.

The comparison table

DetectorParametersF1PrecisionRecallReal FPR
ArtifactNet v9.4 ONNX4.2M0.9520.93297.3%13.8%
AI-Music-Detection AST-60s90.8M0.8400.84883.1%28.9%
CLAM (MoM)194.3M0.7870.71188.3%69.7%
SpecTTTra α-120s18.7M0.7770.88069.5%18.4%
Deezer ISMIR fakeprint LR3.6K0.7540.90664.6%13.0%
FST (Mippia)174.4M0.7350.98458.7%1.8%
DeepFense EAT+Nes2Net0.6500.58972.4%97.8%
SpecTTTra β-5s18.7M0.5630.88441.3%10.5%

Why you should care about false positives

A label, distributor or moderator reviewing a catalog does not only pay for missed AI tracks. Every genuine release incorrectly flagged consumes review time and can damage a creator's standing. In this run two systems with seemingly useful recall flagged more than two-thirds of real music — recall alone is a trap. FST made the opposite trade: the fewest false alarms, but it missed far more AI tracks.

The right tool is the one whose error profile matches your volume and tolerance. If you review thousands of tracks and one false positive is a real fight, a low-FPR tool wins. If you need to catch high volumes of AI music and can spend human time on triage, recall-and-precision balance matters more.

Who each tool suits

Best F1 on this test

ArtifactNet

Highest F1 with 97.3% recall at a manageable 13.8% FPR, and it is the only lightweight model here (4.2M params) with a free in-browser demo. Good default if you want the strongest overall balance. Try it free: demo.intrect.io.

Best precision of the set

FST (Mippia)

1.8% false-positive rate on real music — the safest choice when a wrong flag is expensive. You trade recall (58.7%): it will let more AI tracks through than the leaders.

High confidence, low volume

AI-Music-Detection AST-60s

Solid precision (0.848) and recall (83.1%). A reasonable middle option if ArtifactNet's FPR is more than your operation tolerates.

Segment-level signals

CLAM / SpecTTTra

Useful research references. CLAM catches AI sections but flags ~70% of real music in this run; SpecTTTra trades recall for a lower false-alarm rate. Both are heavier.

What this comparison does not prove

What detector can't do alone

Detecting AI music answers where a track may have come from. Fixing the audible artifacts it leaves on the audio is a separate job — this is what the de-artifact plug-in and online cleanup handle by removing RVQ ghosting and codec residue.

Sources and revision policy

Primary sources: the 8-way field report, the ArtifactBench dataset card, the public result table, and the ArtifactNet paper. This is a dated snapshot: a rerun gets a new report or an explicit correction, never a silent overwrite.

Try the top-F1 detector on your own file

Upload a track and get an AI-vs-human verdict with forensic evidence — free, in the browser, no account, nothing retained.

Run a free AI-music detection analysis

First analysis is free. Treat the result as forensic evidence to review, not an automatic verdict.

Run a free analysis →