In plain words: Ingot checks whether an AI was tested on questions it had already seen — and shows you the words.
Ingot is independent verification for the AI training data supply chain: is this benchmark inside this corpus — answered with matching text, not a score you take on faith.
A shared 10-gram — ten identical words in a row — between a benchmark and a corpus is a fact you can display side by side. Authorship is a statistical claim about a distribution. Both are worth building; only one of them can be handed to a skeptic who then agrees.
Index the benchmark, stream the corpus once, show every match with its surrounding text. 7.0 MB/sec end to end on real gzipped corpus shards, in the browser or on the command line.
Human-authored or machine-generated, measured against named reference corpora. Honest detection floor is 50% contamination, and the four reasons are known.
Which benchmarks appear in which public corpora, reproducible from the same files, growing more discriminating with every corpus added.
Every number here is public. None of it depends on anyone's résumé.
Quality and neutrality now outrank price when labs choose a vendor. So vendors began publishing provenance paperwork — about their own labour. Self-certification survives exactly as long as nobody's money is at risk.
In plain words: does delivered data read like it was written by the humans it was billed as? Structural statistics measured against two named reference corpora, one human-authored and one machine-generated. No model call, no network, deterministic and reproducible. This is the second product, and the numbers below are the honest limits of what it can do today.
Schema mismatches, unparseable lines and zero-variance batches are reported, never silently dropped.
Sentence-length variation, reused openings and closings, cluster tightness, duplicate rate, author fingerprints.
Each signal is weighted by its measured signal-to-noise, and declines when it sits outside its calibrated range.
A purity mark, per-signal contribution, flagged records, and every input stamped so the result can be reproduced.
Human-authored corpus contaminated with machine-generated records at known ratios. Batch size 1,200. Five draws per level. References calibrated on one random half, every scored record from the held-out half.
First level landing more than two control standard deviations below clean. 25% is marginal at about 1.8σ.
One of eight clean human batches would be flagged at the operating threshold. Reported, not buried.
Inside the noise on this corpus pair. Claiming otherwise would be the one unrecoverable error.
The reference pair isn't prompt-matched and the machine side is 2023-era. Regenerating identical prompts with a current model removes both problems, and that's the next run.
Each of these was a bug first, caught by running the experiment and reading the output instead of trusting it.
Is the benchmark inside the corpus, and was the delivered data produced by the human it was billed as. Run either before payment, or before training.
Red-team a reward function or rubric — the scoring rules a model is trained against — to find what scores high while failing the task. Same disease, one layer up.
Third-party AI Act documentation a regulator accepts precisely because it isn't self-issued.
Mercor will not certify Mercor. Anthropic will not ship a product that audits Anthropic's own training data. The referee cannot play for a team — which is exactly why this is a company and not a feature.