Intrect / Artifact Observatory / Report 001
Field report 001

Eight Public AI-Music Detectors on One Test Set

Eight detectors, the same 2,104 audio files and their public decision thresholds. The useful result is not only who ranks first; it is how differently the systems trade missed AI tracks against real music falsely flagged as AI.

Published September 14, 2026Intrect ResearchArtifactBench v1.1

Result in one line: ArtifactNet v9.4 had the highest F1 in this run, while FST had the lowest false-positive rate. Those are different operating objectives; neither number makes a detector an automatic proof of authorship.

The controlled comparison

The run used the same restored ArtifactBench v1.1 files for every adapter: 1,388 AI-generated tracks and 716 real tracks. Scores were evaluated at each public adapter's fixed threshold of 0.5. Parameter counts, preprocessing adapters and per-source results are preserved in the public result record.

DetectorParametersF1PrecisionRecallReal FPR
ArtifactNet v9.4 ONNX4.2M0.9520.93297.3%13.8%
AI-Music-Detection AST-60s90.8M0.8400.84883.1%28.9%
CLAM (MoM)194.3M0.7870.71188.3%69.7%
SpecTTTra α-120s18.7M0.7770.88069.5%18.4%
Deezer ISMIR fakeprint LR3.6K0.7540.90664.6%13.0%
FST (Mippia)174.4M0.7350.98458.7%1.8%
DeepFense EAT+Nes2Net0.6500.58972.4%97.8%
SpecTTTra β-5s18.7M0.5630.88441.3%10.5%

Why the false-positive column matters

A label or distributor reviewing a catalog does not pay only for missed AI tracks. Every genuine release incorrectly escalated consumes review time and can harm a creator. In this run, two systems with superficially useful recall flagged more than two thirds of the real partition. FST made the opposite trade: the fewest false alarms, but missed more AI tracks. The correct threshold therefore depends on whether the detector is a triage signal, an investigative aid or one input to a human appeal process.

What this report does not prove

A separate task: removing audible artifacts

Detection asks where a track may have come from. Restoration asks whether codec-like residue can be reduced without damaging the music. The two clips below demonstrate the second task only. They are the same public, curated Suno example used on the de-artifact product page; they are not extra detector test rows.

Original AI mix

Suno 032 · dry input

Cleaned mix

Same excerpt · de-artifact AI preset

Sources and revision policy

Primary sources: ArtifactBench dataset card, the 8-way result table, and the ArtifactNet paper. This is a dated snapshot. A changed partition or rerun will receive a new report or an explicit correction note instead of silently replacing these figures.

Test the maintained detector on your own file

The first analysis is free. Treat the result as forensic evidence to review, not a verdict to enforce automatically.

Run a free analysis →