evalseal
ASSAY OFFICE FOR AI EVALS

Independent third party · since 2026

“Trust us”
isn't proof

Nearly every AI capability claim is issued by the vendor that benefits from it, and almost none can be reproduced. EvalSeal freezes one evaluation run into a publicly verifiable seal — dataset hash, model build, random seed, every trace, all signed. Anyone can re-run it, compare it, or overturn it.

Sealed Evaluation Record Ed25519 · Reproducible 7F2A·93C1 claude-opus-5

Marks struck on the seal — open any one

Dataset

SWE-bench Verified · 500 tasks · sha256:a3f91c7d…

The task set is frozen and hashed before the run. Swap or edit a single task and the hash stops matching.

Why scores
don't hold up

All four happen daily, and none of them require anyone to cheat. The default evaluation process simply does not produce evidence.

Contamination

The test set is in the training set

Public benchmarks start leaking the day they ship. Scores climb month over month while real capability stays flat. Without a frozen hash and a release date, you cannot tell whether the model improved or memorized.

Irreproducible

The same suite gives a different number

Seed, temperature, prompt version, tool set — miss any one and the second run is a different experiment. Most evaluation reports don't even have a field to record them in.

Drift

The version in the report isn't the one you call

The report says the June build. The endpoint has swapped weights twice since. The model name never changed, the behaviour did, and the PDF in your hand will never update.

Single party

The referee is also the contestant

Vendors run their own evals, pick their own subsets, write their own reports. Not necessarily dishonest — but structurally impossible to refute. Anything that cannot be refuted is not evidence.

One seal in full

SEAL 7F2A-93C1-E0B4 · publicly verifiable
RunSWE-bench Verified · 500 tasks · single pass
Subjectclaude-opus-5 @ 2026-06-11 · build 4e19a2
Datasetsha256:a3f91c7d0e88b2 frozen
Conditionsseed=1729 · temperature=0.0 · top_p=1.0
Result341 / 500 passed · 68.2% (±0.0, deterministic)
Traces500 records · 12.4 MB · sha256:7e21b0f4c95a17
Environmentisolated container · no network egress · image sha256:b81f22de
Sealed at2026-07-28T09:14:02Z
Signatureed25519:9c4f8b1e6a03d7f2… valid

How a seal is struck

01 / Intake

Submit

Point us at your eval suite and dataset. Both are hashed and frozen. Any change after that produces a new seal number — the old seal is never overwritten.

02 / Testing

Assay

The run executes in an isolated container with no network egress, recording every call's input and output. The runtime image is hashed too.

03 / Marking

Hallmark

The seal is issued, signed, and given a public address. Your customers don't have to trust you, and they don't have to trust us — they can re-run it themselves.

A conclusion that can be overturned is the only kind worth believing. EvalSeal exists to make overturning it possible.