The test set is in the training set
Public benchmarks start leaking the day they ship. Scores climb month over month while real capability stays flat. Without a frozen hash and a release date, you cannot tell whether the model improved or memorized.
Independent third party · since 2026
Nearly every AI capability claim is issued by the vendor that benefits from it, and almost none can be reproduced. EvalSeal freezes one evaluation run into a publicly verifiable seal — dataset hash, model build, random seed, every trace, all signed. Anyone can re-run it, compare it, or overturn it.
Marks struck on the seal — open any one
Dataset
SWE-bench Verified · 500 tasks · sha256:a3f91c7d…
The task set is frozen and hashed before the run. Swap or edit a single task and the hash stops matching.
All four happen daily, and none of them require anyone to cheat. The default evaluation process simply does not produce evidence.
Public benchmarks start leaking the day they ship. Scores climb month over month while real capability stays flat. Without a frozen hash and a release date, you cannot tell whether the model improved or memorized.
Seed, temperature, prompt version, tool set — miss any one and the second run is a different experiment. Most evaluation reports don't even have a field to record them in.
The report says the June build. The endpoint has swapped weights twice since. The model name never changed, the behaviour did, and the PDF in your hand will never update.
Vendors run their own evals, pick their own subsets, write their own reports. Not necessarily dishonest — but structurally impossible to refute. Anything that cannot be refuted is not evidence.
Point us at your eval suite and dataset. Both are hashed and frozen. Any change after that produces a new seal number — the old seal is never overwritten.
The run executes in an isolated container with no network egress, recording every call's input and output. The runtime image is hashed too.
The seal is issued, signed, and given a public address. Your customers don't have to trust you, and they don't have to trust us — they can re-run it themselves.
A conclusion that can be overturned is the only kind worth believing. EvalSeal exists to make overturning it possible.