Why this exists

Every model quotes a benchmark.
Nobody can show it wasn't in the data.

In plain words: Ingot checks whether an AI was tested on questions it had already seen — and shows you the words.

Ingot is independent verification for the AI training data supply chain: is this benchmark inside this corpus — answered with matching text, not a score you take on faith.

runs on your machine · no upload · no retention · no API key
The product today

Contamination leads because it is not an inference

A shared 10-gram — ten identical words in a row — between a benchmark and a corpus is a fact you can display side by side. Authorship is a statistical claim about a distribution. Both are worth building; only one of them can be handed to a skeptic who then agrees.

SHIPPED

Contamination scanning

Index the benchmark, stream the corpus once, show every match with its surrounding text. 7.0 MB/sec end to end on real gzipped corpus shards, in the browser or on the command line.

SHIPPED · LIMITED

Batch provenance

Human-authored or machine-generated, measured against named reference corpora. Honest detection floor is 50% contamination, and the four reasons are known.

NEXT

The public registry

Which benchmarks appear in which public corpora, reproducible from the same files, growing more discriminating with every corpus added.

The hole

Nine figures of spend, certified by the seller

Every number here is public. None of it depends on anyone's résumé.

$1B+discussed by a single lab for RL environments in one year — The Information, September 2025
$85–200/hrthe going rate for credentialed expert data — Mercor's published expert rates
Aug 2026the EU AI Act's high-risk obligations, documented provenance included, begin applying — Regulation (EU) 2024/1689

Quality and neutrality now outrank price when labs choose a vendor. So vendors began publishing provenance paperwork — about their own labour. Self-certification survives exactly as long as nobody's money is at risk.

How the provenance scanner works

Six signals, none of them a vibe check

In plain words: does delivered data read like it was written by the humans it was billed as? Structural statistics measured against two named reference corpora, one human-authored and one machine-generated. No model call, no network, deterministic and reproducible. This is the second product, and the numbers below are the honest limits of what it can do today.

01 · LOAD

Refuse, don't guess

Schema mismatches, unparseable lines and zero-variance batches are reported, never silently dropped.

02 · MEASURE

Burstiness & scaffolding

Sentence-length variation, reused openings and closings, cluster tightness, duplicate rate, author fingerprints.

03 · COMPARE

Against distributions

Each signal is weighted by its measured signal-to-noise, and declines when it sits outside its calibrated range.

04 · REPORT

Evidence, not a verdict

A purity mark, per-signal contribution, flagged records, and every input stamped so the result can be reproduced.

Measured, not promised

What it detects today — including what it can't

Human-authored corpus contaminated with machine-generated records at known ratios. Batch size 1,200. Five draws per level. References calibrated on one random half, every scored record from the held-out half.

100500 machine contamination purity clean human control 89.9 ± 4.7 0% 5% 10% 25% 50% 88.0 90.2 87.6 81.2 56.4
spiked batches, mean ± sd clean human control band
Detection floor: 50%

First level landing more than two control standard deviations below clean. 25% is marginal at about 1.8σ.

False positives: 13%

One of eight clean human batches would be flagged at the operating threshold. Reported, not buried.

5% and 10%: not detectable

Inside the noise on this corpus pair. Claiming otherwise would be the one unrecoverable error.

Why, and what's next

The reference pair isn't prompt-matched and the machine side is 2023-era. Regenerating identical prompts with a current model removes both problems, and that's the next run.

Discipline

What Ingot refuses to do

Each of these was a bug first, caught by running the experiment and reading the output instead of trusting it.

Where this goes

One company, three acts

Now · built

Contamination & provenance

Is the benchmark inside the corpus, and was the delivered data produced by the human it was billed as. Run either before payment, or before training.

Act II

Reward-signal integrity

Red-team a reward function or rubric — the scoring rules a model is trained against — to find what scores high while failing the task. Same disease, one layer up.

Act III

Provenance of record

Third-party AI Act documentation a regulator accepts precisely because it isn't self-issued.

Mercor will not certify Mercor. Anthropic will not ship a product that audits Anthropic's own training data. The referee cannot play for a team — which is exactly why this is a company and not a feature.