When to use: Your benchmark has versions (HumanEval got
v2, MMLU got de-contaminated, GSM8K had a leak fixed). You want to pin which one you committed against.
Two papers can both report "62% on HumanEval" against two different dataset revisions and you cannot tell which is which.
dataset:
id: "humaneval"
hash: "7c33e0a4b2d1f8e6c5a4938271605f4e3d2c1b0a99887766554433221100ffee" # SHA-256 of the canonical bytes (64 hex)dataset is the name (human-readable); dataset_hash is the content commitment (machine-verifiable). Pre-compute the hash once and paste it.
For a single file:
sha256sum humaneval.jsonl
# 7c33e0a4b2... humaneval.jsonlFor a directory or multi-file dataset, you need a canonicalisation step so anyone re-deriving gets the same hash:
# Sort filenames, concat, hash
( cd dataset_dir && find . -type f | sort | xargs cat ) | sha256sumThis is fragile. Better: tar the directory deterministically:
tar --sort=name \
--owner=0 --group=0 --numeric-owner \
-cf - dataset_dir/ | sha256sumBetter still: if your benchmark is published on HuggingFace Datasets, use the platform's content SHA:
from datasets import load_dataset
ds = load_dataset("openai/human-eval", split="test")
print(ds.info.dataset_size, ds.info.download_size)
# the dataset card on HF includes a `sha` field that pins the revisionYour manifest's dataset_hash then references that platform SHA explicitly:
dataset:
id: "openai/human-eval @ hf-revision d8f3e1a2"
hash: "d8f3e1a2c4b6098e7f5a3c1b9d8e7f6a5c4b3d2e1f0a9b8c7d6e5f4a3b2c1d0e" # SHA-256 of that revisionWe use huggingface: instead of sha256: here because the underlying hash is the platform's, not yours. The point is: anyone who reads your manifest must be able to fetch exactly the dataset you committed against.
1. The dataset gets updated and you don't notice. HumanEval v1 and v2 differ in 12 problem statements. If you committed dataset: "humaneval" without a hash, no one knows which revision your "62%" refers to. Always commit a hash; never commit just a name.
2. You hash the un-canonicalised version. You compute the hash of humaneval.jsonl after your local pre-processing (lowercase, whitespace strip). Your colleague re-derives against the original file and gets a different hash. Pin: hash the canonical published bytes, never your local working copy.
3. The dataset is private. If dataset_hash is sha256:abc... of a private file, no one can verify. Either:
- Publish the file (and live with the consequences)
- Publish a cryptographic commitment to the file structure (Merkle root of N samples) and keep the file private
- Accept the claim is internally meaningful but externally unverifiable
The third is OK for internal audits; not OK for public claims.
4. Multi-language tokenisation drift. A benchmark file with Unicode strings can produce different bytes depending on normalisation form (NFC vs NFD). Force NFC before hashing:
iconv -f UTF-8 -t UTF-8 humaneval.jsonl | uconv -f UTF-8 -t UTF-8 -x "::NFC;" | sha256sum-
Pinning by URL. URLs are not content-stable. The maintainers can update the file without changing the URL.
-
Pinning by paper version. Papers reference datasets but don't carry their bytes. You need the bytes.
-
Hashing the schema (column names) instead of the content. Schema-only hashes don't catch row-level changes. Hash everything.
dataset_hash makes contamination detectable, not prevented. A determined publisher can hash a clean eval set, run on a contaminated one, and report the clean hash. PRML §8.1 names this. v0.2 P-02 (runner_attestation) addresses it; v0.1 does not.
What dataset_hash does prevent: a publisher claiming they used dataset A when they actually ran on a different file they call "A". The content hash makes that lie immediately falsifiable.
For CI-level verification: see Pattern 5 — CI gate.