Provir

A lossless container is not a lossless file. Provir measures which one you have.

FLAC, WAV, AIFF and ALAC can all carry audio that was lossy earlier in its life. Provir decides which, from physical measurements — spectral geometry, quantisation-lattice structure, temporal stationarity of high-frequency content, re-encode fixed-point distance — under a rule that a conviction requires agreement from two independent evidence families.

The detector is not the interesting part. Six claim groups below are re-derived from their raw ledgers by an automated harness before they can be quoted, and the harness reports disagreement — including disagreement with its author. When it was first wired it found three published claims wrong. All three had been wrong silently.

Corrected 2026-08-20. This read “every published number below is re-derived”, and that was not true of this page. The harness asserts 14 field values across 5 claim groups; this page carries 33 bolded figures, of which 11 carry a value the harness checks. Rows without a re-derived marker are pinned or prose-only, and are now labelled as such rather than covered by a blanket sentence. The overclaim was found by an audit of our own repository and is corrected in place rather than quietly reworded — a page whose opening promise is verification is the last place a reader should have to check the promise.

Claims harness last run 19 August 2026
5 re-derived and reproducing 1 under re-measurement 1 hash-pinned only, not re-run
Measured results

The ledger

Each row carries its instrument, its independent sample size, and its bound. Where a figure has a caveat that weakens it, the caveat is part of the claim rather than a footnote.

With its neural component ablated, the chain still reaches 97.4% review-tier recall
re-derived
On a frozen held-out corpus, at zero false positives across its genuine sources. Review tier means flagged for a human, not convicted. Within that chain the quantisation-lattice measurement — implemented from published literature, with no neural references — is worth 3.4 percentage points.
recall 97.4% without lattice 94.0% genuine sources flagged 0 / 32 hard conviction 4.7%
The model's net effect on conviction is to withhold two, not to add any — 4.7% with the full chain against 4.9% without it. The corpus is clean single-generation transcodes, so it prices the easier half of the threat model.
Structural detection discriminates against 28 near-identical decoys
re-derived
Across a lawful corpus of 620 independent recordings, 28 carry a spectral wall in the same 21.3–21.8 kHz band as a proven transcode. The geometry probe fires on none of them; the transcode does fire. The separating mechanism is contrast-and-void geometry above the wall — not the wall's frequency or width, which an earlier draft credited and our own audit refuted.
recordings 620 releases 56 false fires 0 decoys in band 28
We deliberately do not lead with "zero false fires". The probe abstains before reaching its decision rule on all 620, by its own contract — so the decided subset is empty and any false-fire rate over it is 0/0. That is not a number and we do not quote one. A previously published bound over the handful that decided was struck rather than restated.
MP3 detection survives nine laundering attacks; the AAC arm does not
re-derived
Every laundered MP3 in the battery is still caught, with no conviction of a control. The honest counterpart is stated in the same breath: a 2% DJ pitch shift defeats the AAC arm entirely. The asymmetry stated plainly is worth more than the win alone.
laundered files 144 threat caught 72 / 72 convictions of controls 0 vectors 7
Conviction capability on a frozen corpus, with genuine sources untouched
re-derived
Hard conviction on constructed transcodes, against zero convictions of the genuine sources those transcodes were built from. Every conviction is flag-named and independently re-derivable.
fakes convicted 372 / 768 sources convicted 0 / 32
Independent n is 32 recordings, never 768 files. Constructed transcodes are the easy half; the wild-provenance tiers are quoted separately and never blended with these.
A component that fails its own bar does not appear in our claims
re-derived
The neural classifier was retrained on a provenance-clean corpus and lost a held-out comparison to the model it was meant to replace, at matched zero-false-positive thresholds. It is barred from every published claim, and no neural component can reach a conviction. The cause was measured as training-class starvation rather than contamination.

Corrected 2026-08-20. This read “the shipping detection chain contains no neural component at all”, and that was false: two small ONNX classifiers do load and run on every scan. The true and checkable version is narrower — neither is read by any conviction branch, and their only verdict effect is to move a file from GENUINE into review. Recorded rather than quietly reworded, because a claim about our own architecture that a reader could disprove by cloning the repository is the kind this project exists to object to.
retrained 82.6% incumbent 87.7% McNemar p = 0.0029 files scored 743
Independent replication of one mechanism, by the author of a competing tool
his measurement, not ours
A high-frequency temporal-seam measurement of ours was implemented from a four-line description by the author of the leading open-source tool in this problem space, and scored on his corpus with his labels. It reached AUC 0.84 and 0.89 on the two arms where his own detector families are blind — and his best arm is one where ours is near-blind, which is a better argument for holding both than either of us set out to make.

Status corrected 2026-08-20. This row was marked “pinned, not re-run”, which claimed a harness status it never had: our harness has no group for this claim, and the figures appear in neither the freeze file nor the draft. It is neither hash-pinned nor prose-pinned by us — it is a third party's measurement on his corpus with his labels, which is exactly what makes it worth showing and exactly why we cannot re-derive it. Attributing a verification status to a row our own harness has never seen was the wrong kind of mistake for this page to make.
Opus AUC 0.84 MP3-320 AUC 0.89 our weakest arm 0.47
This row is hash-pinned but not verified by re-running, and the harness says so rather than reporting it green. A hash proves an artefact has not moved. It says nothing about whether the sentence quoting it was ever true — which is exactly how one of our own claims stayed wrong for four days while its pin held.
Detection of AI-generated music, as a separate question
registered, not wired
Generated audio carries its own measurable signature, and we have one — a conviction rung with a named mechanism, pre-registered before it was measured. It is deliberately not switched on. Its preconditions are published and unmet: a stated minimum of lawful, sparse and producer-processed material must be in the corpus before it may carry a verdict.
And the honest limit is in the mechanism itself. What this measures is the fingerprint of an export, not of a machine having composed something. Those two questions get conflated constantly and we will not conflate them: a generated track re-recorded through an analogue chain is a different problem, and this does not solve it.
False-conviction rate on lawful recordings
under re-measurement
Withheld. The corpus this figure was cut from has been superseded, and the number is being re-derived against the rebuilt bench before it is quoted again. It will return with its dependencies stated.
Kept, not deleted

The graveyard

Approaches that failed are recorded with the measurement that killed them, in the same ledger as the ones that worked. An approach with no gravestone was never really tested.

Pre-echo as a corroborating tell
refuted
Noise. It cannot corroborate anything, and the result was confirmed against a second implementation rather than only our own.
AUC 0.517AUC 0.575
A stereo rule for MP3-320 detection
retired
Retired after we established that a 20 kHz lowpass — an ordinary mastering choice — manufactured the overwhelming majority of its own convictions.
convictions manufactured 92%
A 20 kHz brickwall flag as active evidence
retired
Priced honestly and removed: it produced false flags in quantity and caught nothing the chain did not already have. Retiring it cost zero recall.
false flags 46unique catches 0
A frame-grid periodicity readout
refuted
It measured phase, not period. The statistic was real; what we believed it meant was not.
A level-guarded normalisation inside our own dead-band statistic
removed
Our own defect, found by our own review rather than by an adversary. The guard made the statistic peak-relative below its threshold and absolute above it — two different statistics with a seam between them, where an inaudible level difference moved the reading across a decision boundary.
level change 0.002 dBreading moved 2.8×
Why the graveyard is on the front page. A detector that only publishes its successes cannot be audited, and in a forensic context an unauditable accuser is as dangerous as the fraud it claims to find. The failures above are the evidence that the successes were tested the same way.
Around the detector

It is an application, not a script

The measurement work sits inside a desktop tool built for people who own music, not for people who own a terminal. Two things there are worth naming on their own.

A free desktop application, not a command line
shipped
Scanning, spectrogram inspection, library organisation and the verdicts themselves, in an interface anyone can use. It is released free, under the same licence as the engine, and it is not part of what this project is asking to be funded — it already exists.
The tempo octave problem, solved
shipped
Beat detectors habitually report half or double the real tempo, because for drum & bass, hardcore and hard trance the felt backbeat is half the production tempo. Every major DJ tool trips on it. A genre-aware octave corrector fixes it: zero detector-wrong across 494 tracks with embedded tag ground truth, and 30 cases where the detector was right and the tag was wrong.
detector wrong 0 / 494 beats the tag 30 controlled octave fix 17 / 17 library BPMs repaired 1,100
The raw match rate is 80% and we do not quote it, because the ground truth is itself half-tagged — the disagreements are mostly the tags being wrong. That is also the honest reason the headline is "zero errors", not an accuracy percentage. A small number of non-octave corrections remain plausible but unverified by ear, and are recorded as such.
Method

What holds the numbers up

Four rules, each of which has cost us a published figure at least once.

A conviction needs two independent evidence families
Not two readings of one measurement. Distinguishing those two is harder than it sounds — we have found and fixed eight separate places where something that could not stand alone was being accepted as the second family.
Abstention is a verdict
Where the evidence does not reach, the answer is we cannot tell, recorded as a named abstention with its reason. A statistic that returns nothing must never be read as returning "clean".
Every figure travels with its denominator and its bound
Independent sample size counts recordings and releases, never files. A false-positive rate is quoted beside the detection rate that earned it, in the same sentence.
Corrections may only ever push conservative
A change that makes our own numbers look better needs a higher bar than one that makes them look worse. Where a cleaner denominator would widen a bound, we take the wider bound.
The machine in the room

Built with AI, and that is checkable rather than claimed

Expert opinion produced by feeding data to a model and accepting what comes out is a real and growing problem, and courts are starting to ask for the prompts. The distinction that matters is not whether a machine was involved. It is that these results are empirical, not hypothetical — measured on named corpora, with stated denominators and bounds, and with the failures published beside the successes. A model's output accepted on faith is a hypothesis wearing a conclusion's clothes. Everything below can be re-measured by someone who does not trust us.

A verdict here is a measurement, not an opinion
reproducible
No neural component reaches a conviction, and no language model is anywhere in the path. A verdict is deterministic code run against a file: spectral geometry, lattice structure, temporal statistics. Run it again and you get the same answer. Two small ONNX classifiers do run, and either can move a file into the review tier — neither can convict, and with both ablated the chain still reaches 97.4% review-tier recall. There is no prompt and nothing generative between the audio and the result — so there is nothing to subpoena, because there is nothing to reconstruct. (Corrected 2026-08-20: this previously claimed no neural component at all.)
AI built the instrument; it does not operate it
disclosed per change
Generative AI was used throughout development, heavily and continuously. That is disclosed at the granularity of the individual change: model and version attached to each commit, in the published history, timestamped — not a retrospective statement written once for a form. Session transcripts are retained and available on request.
commits carrying model attribution 874 of 1,433
The human judgements are logged, including the ones that overruled the model
recorded
Every decision only a person could make — a listening judgement, a ruling that moves a denominator, a disclosure call — is recorded with what it overruled, its consequence, and where the evidence sits. That deliberately includes the occasions where the model proposed something and was wrong: on one recent day it offered a piece of supporting evidence that was structurally impossible, and the objection took a domain expert one sentence.
The corpus and its adjudications are not model-generated and could not be. Physical discs, receipted purchases, and a listening judgement per file recorded with its basis and its date — several of them made years before this instrument existed, which is what makes them evidence rather than the tool grading its own homework.
The test we would want applied to anyone. Not "was a machine involved" — it was, and pretending otherwise is the actual red flag. The question is whether the conclusion survives without trusting the person who produced it. Here the answer is a rerun: the code, the ledgers and the harness are released together, and every figure above is re-derived from raw data by a program that reports against its own author.