Each row carries its instrument, its independent sample size, and its bound. Where a figure has a
caveat that weakens it, the caveat is part of the claim rather than a footnote.
With its neural component ablated, the chain still reaches 97.4% review-tier recall
re-derived
On a frozen held-out corpus, at zero false positives across its genuine sources. Review tier means
flagged for a human, not convicted. Within that chain the quantisation-lattice measurement —
implemented from published literature, with no neural references — is worth 3.4 percentage points.
recall 97.4%
without lattice 94.0%
genuine sources flagged 0 / 32
hard conviction 4.7%
The model's net effect on conviction is to withhold two, not to add any — 4.7% with the full
chain against 4.9% without it. The corpus is clean single-generation transcodes, so it prices the
easier half of the threat model.
Structural detection discriminates against 28 near-identical decoys
re-derived
Across a lawful corpus of 620 independent recordings, 28 carry a spectral wall in the same
21.3–21.8 kHz band as a proven transcode. The geometry probe fires on none of them; the
transcode does fire. The separating mechanism is contrast-and-void geometry above the wall — not
the wall's frequency or width, which an earlier draft credited and our own audit refuted.
recordings 620
releases 56
false fires 0
decoys in band 28
We deliberately do not lead with "zero false fires". The probe abstains before reaching its
decision rule on all 620, by its own contract — so the decided subset is empty and any false-fire
rate over it is 0/0. That is not a number and we do not quote one. A previously published bound
over the handful that decided was struck rather than restated.
MP3 detection survives nine laundering attacks; the AAC arm does not
re-derived
Every laundered MP3 in the battery is still caught, with no conviction of a control. The honest
counterpart is stated in the same breath: a 2% DJ pitch shift defeats the AAC arm entirely.
The asymmetry stated plainly is worth more than the win alone.
laundered files 144
threat caught 72 / 72
convictions of controls 0
vectors 7
Conviction capability on a frozen corpus, with genuine sources untouched
re-derived
Hard conviction on constructed transcodes, against zero convictions of the genuine sources those
transcodes were built from. Every conviction is flag-named and independently re-derivable.
fakes convicted 372 / 768
sources convicted 0 / 32
Independent n is 32 recordings, never 768 files. Constructed transcodes are the easy
half; the wild-provenance tiers are quoted separately and never blended with these.
A component that fails its own bar does not appear in our claims
re-derived
The neural classifier was retrained on a provenance-clean corpus and lost a held-out
comparison to the model it was meant to replace, at matched zero-false-positive thresholds. It is
barred from every published claim, and no neural component can reach a conviction. The
cause was measured as training-class starvation rather than contamination.
Corrected 2026-08-20. This read “the shipping detection chain contains no
neural component at all”, and that was false: two small ONNX classifiers do load and run on
every scan. The true and checkable version is narrower — neither is read by any conviction
branch, and their only verdict effect is to move a file from GENUINE into review. Recorded rather
than quietly reworded, because a claim about our own architecture that a reader could disprove by
cloning the repository is the kind this project exists to object to.
retrained 82.6%
incumbent 87.7%
McNemar p = 0.0029
files scored 743
Independent replication of one mechanism, by the author of a competing tool
his measurement, not ours
A high-frequency temporal-seam measurement of ours was implemented from a four-line description by
the author of the leading open-source tool in this problem space, and scored on his corpus
with his labels. It reached AUC 0.84 and 0.89 on the two arms where his own detector
families are blind — and his best arm is one where ours is near-blind, which is a better argument
for holding both than either of us set out to make.
Status corrected 2026-08-20. This row was marked “pinned, not re-run”,
which claimed a harness status it never had: our harness has no group for this claim, and the
figures appear in neither the freeze file nor the draft. It is neither hash-pinned nor prose-pinned
by us — it is a third party's measurement on his corpus with his labels, which is exactly
what makes it worth showing and exactly why we cannot re-derive it. Attributing a verification
status to a row our own harness has never seen was the wrong kind of mistake for this page to make.
Opus AUC 0.84
MP3-320 AUC 0.89
our weakest arm 0.47
This row is hash-pinned but not verified by re-running, and the harness says so rather than
reporting it green. A hash proves an artefact has not moved. It says nothing about whether the
sentence quoting it was ever true — which is exactly how one of our own claims stayed wrong for
four days while its pin held.
Detection of AI-generated music, as a separate question
registered, not wired
Generated audio carries its own measurable signature, and we have one — a conviction rung with a
named mechanism, pre-registered before it was measured. It is deliberately not switched on.
Its preconditions are published and unmet: a stated minimum of lawful, sparse and
producer-processed material must be in the corpus before it may carry a verdict.
And the honest limit is in the mechanism itself. What this measures is the fingerprint of an
export, not of a machine having composed something. Those two questions get conflated
constantly and we will not conflate them: a generated track re-recorded through an analogue chain
is a different problem, and this does not solve it.
False-conviction rate on lawful recordings
under re-measurement
Withheld. The corpus this figure was cut from has been superseded, and the number is being
re-derived against the rebuilt bench before it is quoted again. It will return with its
dependencies stated.