SanctionsKit AI sanctions study: complete evidence archive for the original v1 experiment

This is an expanded evidence package, not a new experiment or a change to the results. The original experiment ran on 8 October 2026 and contains 96 selected company identities and 1,152 evaluated requests. Follow-up experiments, if any, are separate. The original v1 data, exact requests, visible model outputs, citations, scoring records and analysis code are unchanged. The original README is retained as README-v1-original.txt.

What is added
- Complete archived authoritative SDN XML used for the reference labels (19,416 records), not just company excerpts. It includes all record types because absence checks require the whole list.
- Original May, July and October deletion notices, plus the two earlier XML snapshots needed by the original selection code.
- Original selection-code dependencies, internal source/review records, and an offline verifier for all 96 reference labels, original cohort/frame selection, all 1,152 exact requests, and all seven published scoring/comparison groups.
- A complete file manifest with SHA-256 hashes and a disclosed deterministic lookup baseline on the same selected panel.
- A later cost-accounting correction, its original and revised receipts, reconstruction code, and two separately labeled unscored engineering captures; no experimental result changes.

The immutable original methodology and README say full government snapshots are available on request. That earlier availability statement is superseded by this complete archive; those historical files have deliberately not been silently edited.

Reproduce without an API key or network
Use Python 3.10 or later with the standard library. From this extracted directory:
python3 archive-additions/verify_complete_archive.py --root . --output verification-local.json --baseline-output lookup-baseline-local.json

The verifier checks every included file against the manifest; parses the full authoritative XML independently; verifies each selected company's original names and public identifiers against every primary/alias name and supported corporate identifier; reconstructs the original selection frame/cohort from the archived sources; checks retained notice excerpts within the SDN deletion sections; verifies exact scheduled request payloads; and reruns the unchanged scoring analyzer. It fails if the recomputed results differ. It uses a temporary directory for selection reconstruction and does not change the preserved records. Its output files and Python cache files are not part of the original archive manifest.

Where to inspect the evidence
- companies-v1.csv: company dataset and original reference labels.
- methods/api-frozen/schedule.json: every exact input, model identifier, setting and request allocation.
- private/api-execution/: 1,152 sanitized captures with every visible model answer, citation, web action, and usage/completion metadata.
- private/api-review-M.json, -W.json and -E.json: all internal final-text scoring records.
- results-v1.json and private/api-response-analysis-preliminary.{csv,jsonl}: unchanged aggregates and case-level outcomes.
- selected-source-evidence/: all 96 source packets actually supplied in the evidence condition.
- scripts/analyze_api_study.py: original offline scoring code.
- archive-additions/official-sources/: original government bytes.
- archive-additions/internal-review/: source-identity checks, near-name dossiers and wrong-answer corroboration.
- transformation-manifest.json: original raw-capture and sanitized-capture hashes.
- archive-additions/file-manifest.json: file sizes and hashes for this complete package.
- archive-additions/cost/: original/revised cost estimates and reproducible cost inputs. Neither is a provider invoice.

Redactions and scope
The original v1 export preserves all visible answer text and citation annotations exactly. It removes credentials, account/provider identifiers, hidden reasoning output or ciphertext, and non-scientific transport metadata. This package adds no hidden reasoning or credentials. The folder name private is retained from the original analyzer's layout; its contents in this archive are reviewed distributable research evidence, not private customer data. The full XML contains publicly released sanctions records, including personal data published by the authority, and is reproduced solely as the historical source used by this study.

Review was internal and unblinded: reviewers could see model/condition, reference labels and source material. No external expert peer review or blinded human scoring is claimed. Code reproduction checks arithmetic, consistency and the declared reference-lookup procedure; it does not turn these selected cases into a representative population sample or establish legal clearance.

The deterministic lookup baseline uses only each neutral question's primary company name and public company identifiers against the archived official list. Selection and the reference labels already used related exact lookups, so agreement is not an independent product accuracy test. Lookup timings exclude XML loading and indexing, depend on the local machine, and do not measure a hosted API. New AI results and production-screening accuracy are not claimed.

Original files retain their historical labels, including candidate, preliminary and internal freeze terminology. All 1,152 evaluated API answers were ultimately reviewed; the published totals are in results-v1.json. This package is not public preregistration. SanctionsKit designed and funded the study and sells sanctions screening software.
