For eval teams

The floor, the signals,
and where this goes.

Ingot's front door is a writer asking whether her sentences are inside a training corpus. This page is the other half: the measured limits of the provenance scanner, and the record the registry is building. Nothing here is a claim without a number behind it.

The public registry → is the findings themselves — three benchmarks against 21.33 GB of C4, and every match inspected.

The hole

Nine figures of spend, certified by the seller

Every number here is public. None of it depends on anyone's résumé.

$1B+discussed by a single lab for RL environments in one year — The Information, September 2025
$85–200/hrthe going rate for credentialed expert data — Mercor's published expert rates
Aug 2026the EU AI Act's high-risk obligations, documented provenance included, begin applying — Regulation (EU) 2024/1689

Quality and neutrality now outrank price when labs choose a vendor. So vendors began publishing provenance paperwork — about their own labour. Self-certification survives exactly as long as nobody's money is at risk.

How the provenance scanner works

Six signals, none of them a vibe check

In plain words: does delivered data read like it was written by the humans it was billed as? Structural statistics measured against two named reference corpora, one human-authored and one machine-generated. No model call, no network, deterministic and reproducible. This is the second product, and the numbers below are the honest limits of what it can do today.

01 · LOAD

Refuse, don't guess

Schema mismatches, unparseable lines and zero-variance batches are reported, never silently dropped.

02 · MEASURE

Burstiness & scaffolding

Sentence-length variation, reused openings and closings, cluster tightness, duplicate rate, author fingerprints.

03 · COMPARE

Against distributions

Each signal is weighted by its measured signal-to-noise, and declines when it sits outside its calibrated range.

04 · REPORT

Evidence, not a verdict

A purity mark, per-signal contribution, flagged records, and every input stamped so the result can be reproduced.

Measured, not promised

What it detects today — including what it can't

Human-authored corpus contaminated with machine-generated records at known ratios. Batch size 1,200. Five draws per level. References calibrated on one random half, every scored record from the held-out half.

100500 machine contamination purity clean human control 89.9 ± 4.7 0% 5% 10% 25% 50% 88.0 90.2 87.6 81.2 56.4
spiked batches, mean ± sd clean human control band
Detection floor: 50%

First level landing more than two control standard deviations below clean. 25% is marginal at about 1.8σ.

False positives: 13%

One of eight clean human batches would be flagged at the operating threshold. Reported, not buried.

5% and 10%: not detectable

Inside the noise on this corpus pair. Claiming otherwise would be the one unrecoverable error.

Why, and what's next

The reference pair isn't prompt-matched and the machine side is 2023-era. Regenerating identical prompts with a current model removes both problems, and that's the next run.

Where this goes

One company, three acts

Now · built

Contamination & provenance

Is the benchmark inside the corpus, and was the delivered data produced by the human it was billed as. Run either before payment, or before training.

Act II

Reward-signal integrity

Red-team a reward function or rubric — the scoring rules a model is trained against — to find what scores high while failing the task. Same disease, one layer up.

Act III

Provenance of record

Third-party AI Act documentation a regulator accepts precisely because it isn't self-issued.

Mercor will not certify Mercor. Anthropic will not ship a product that audits Anthropic's own training data. The referee cannot play for a team — which is exactly why this is a company and not a feature.