Eight Public AI-Music Detectors on One Test Set
Eight detectors, the same 2,104 audio files and their public decision thresholds. The useful result is not only who ranks first; it is how differently the systems trade missed AI tracks against real music falsely flagged as AI.
Result in one line: ArtifactNet v9.4 had the highest F1 in this run, while FST had the lowest false-positive rate. Those are different operating objectives; neither number makes a detector an automatic proof of authorship.
The controlled comparison
The run used the same restored ArtifactBench v1.1 files for every adapter: 1,388 AI-generated tracks and 716 real tracks. Scores were evaluated at each public adapter's fixed threshold of 0.5. Parameter counts, preprocessing adapters and per-source results are preserved in the public result record.
| Detector | Parameters | F1 | Precision | Recall | Real FPR |
|---|---|---|---|---|---|
| ArtifactNet v9.4 ONNX | 4.2M | 0.952 | 0.932 | 97.3% | 13.8% |
| AI-Music-Detection AST-60s | 90.8M | 0.840 | 0.848 | 83.1% | 28.9% |
| CLAM (MoM) | 194.3M | 0.787 | 0.711 | 88.3% | 69.7% |
| SpecTTTra α-120s | 18.7M | 0.777 | 0.880 | 69.5% | 18.4% |
| Deezer ISMIR fakeprint LR | 3.6K | 0.754 | 0.906 | 64.6% | 13.0% |
| FST (Mippia) | 174.4M | 0.735 | 0.984 | 58.7% | 1.8% |
| DeepFense EAT+Nes2Net | — | 0.650 | 0.589 | 72.4% | 97.8% |
| SpecTTTra β-5s | 18.7M | 0.563 | 0.884 | 41.3% | 10.5% |
Why the false-positive column matters
A label or distributor reviewing a catalog does not pay only for missed AI tracks. Every genuine release incorrectly escalated consumes review time and can harm a creator. In this run, two systems with superficially useful recall flagged more than two thirds of the real partition. FST made the opposite trade: the fewest false alarms, but missed more AI tracks. The correct threshold therefore depends on whether the detector is a triage signal, an investigative aid or one input to a human appeal process.
What this report does not prove
- The restored real subset contained 716 available files, not every row in the purged manifest. The public record explains the restoration boundary.
- A fixed threshold is reproducible, but it is not necessarily the best threshold for a particular catalog or error budget.
- Generator updates and new delivery codecs can change performance after the review date.
- Detection is not provenance. A score should not replace Content Credentials, platform records or a human review and appeal path.
A separate task: removing audible artifacts
Detection asks where a track may have come from. Restoration asks whether codec-like residue can be reduced without damaging the music. The two clips below demonstrate the second task only. They are the same public, curated Suno example used on the de-artifact product page; they are not extra detector test rows.
Original AI mix
Suno 032 · dry input
Cleaned mix
Same excerpt · de-artifact AI preset
Sources and revision policy
Primary sources: ArtifactBench dataset card, the 8-way result table, and the ArtifactNet paper. This is a dated snapshot. A changed partition or rerun will receive a new report or an explicit correction note instead of silently replacing these figures.
Test the maintained detector on your own file
The first analysis is free. Treat the result as forensic evidence to review, not a verdict to enforce automatically.
Run a free analysis →