The floor, the signals,
and where this goes.
Ingot's front door is a writer asking whether her sentences are inside a training corpus. This page is the other half: the measured limits of the provenance scanner, and the record the registry is building. Nothing here is a claim without a number behind it.
The public registry → is the findings themselves — three benchmarks against 21.33 GB of C4, and every match inspected.
Nine figures of spend, certified by the seller
Every number here is public. None of it depends on anyone's résumé.
Quality and neutrality now outrank price when labs choose a vendor. So vendors began publishing provenance paperwork — about their own labour. Self-certification survives exactly as long as nobody's money is at risk.
Six signals, none of them a vibe check
In plain words: does delivered data read like it was written by the humans it was billed as? Structural statistics measured against two named reference corpora, one human-authored and one machine-generated. No model call, no network, deterministic and reproducible. This is the second product, and the numbers below are the honest limits of what it can do today.
Refuse, don't guess
Schema mismatches, unparseable lines and zero-variance batches are reported, never silently dropped.
Burstiness & scaffolding
Sentence-length variation, reused openings and closings, cluster tightness, duplicate rate, author fingerprints.
Against distributions
Each signal is weighted by its measured signal-to-noise, and declines when it sits outside its calibrated range.
Evidence, not a verdict
A purity mark, per-signal contribution, flagged records, and every input stamped so the result can be reproduced.
What it detects today — including what it can't
Human-authored corpus contaminated with machine-generated records at known ratios. Batch size 1,200. Five draws per level. References calibrated on one random half, every scored record from the held-out half.
First level landing more than two control standard deviations below clean. 25% is marginal at about 1.8σ.
One of eight clean human batches would be flagged at the operating threshold. Reported, not buried.
Inside the noise on this corpus pair. Claiming otherwise would be the one unrecoverable error.
The reference pair isn't prompt-matched and the machine side is 2023-era. Regenerating identical prompts with a current model removes both problems, and that's the next run.
One company, three acts
Contamination & provenance
Is the benchmark inside the corpus, and was the delivered data produced by the human it was billed as. Run either before payment, or before training.
Reward-signal integrity
Red-team a reward function or rubric — the scoring rules a model is trained against — to find what scores high while failing the task. Same disease, one layer up.
Provenance of record
Third-party AI Act documentation a regulator accepts precisely because it isn't self-issued.
Mercor will not certify Mercor. Anthropic will not ship a product that audits Anthropic's own training data. The referee cannot play for a team — which is exactly why this is a company and not a feature.