SanctionsKit AI sanctions study — separate follow-up evidence archive v2

This is the follow-up experiment, not a revision of the original v1 results. Its fixed allocation is 96 selected companies, two model identifiers, five conditions and one repetition: 960 scheduled company decisions. Read the full scheduled outcome table, including every missing, operationally failed, truncated and contaminated slot. Seven initially generated answers were preserved; nine documented pre-generation Flex capacity failures led to an availability/accounting amendment. No generated answer was retried or replaced. All initial failures, amended attempts and frozen protocols are retained.

Reproduce locally without API credentials, network or paid requests:
python3 scripts/verify_followup_archive.py --root .

The verifier checks all file hashes and frozen commitments, validates all neutral reference queries against the complete pinned official SDN XML, reproduces recorded deterministic tool results, and reruns the reviewed membership/outcome/cost/latency analysis. Recorded web actions, URLs and visible claims do not expose every retrieved page body or prove which passage the model read. A contradiction between a cited notice and an answer is observable; a retrieval-versus-interpretation causal explanation is not automatically established. Review is internal and unblinded. Arithmetic reproduction is not independent expert peer review, population accuracy, legal clearance or a hosted product benchmark. The separately archived Federal Register PDF 2026-20526 was downloaded during post-answer internal review to corroborate a cited source; its receipt records this timing. It was not supplied to the model as study input and does not revise the frozen reference labels.

The web conditions request maximum tool-call limits of 4 or 8 and low or high search-context settings: W4-low, W8-low and W4-high. These are configured limits, not guaranteed numbers of completed searches. D supplies official document excerpts without answer labels; a historical removal notice does not establish later absence by itself. T lets the model choose exact names and typed corporate identifiers to query the complete pinned official XML. All conditions use medium reasoning and a 4096-token total output allowance; T shares that allowance across its tool turns.

Exact initial inputs/settings are in methods/followup-frozen/schedule.json. The amendment schedule preserves identical original items for the 953 not-yet-generated decisions. All frozen directories are byte-for-byte internal commitments, not public preregistration. The complete original government XML is included in both freezes; source receipts state its original timestamp and hash. The selected panel and reference labels used related exact lookups, so 96/96 deterministic agreement is not an independent production-accuracy test.

private/followup-execution, private/followup-amendment-execution and private/followup-tier-execution hold distributable sanitized evidence, not credentials or customer data. Every visible message, final answer, refusal, citation, web action, function query/result, usage field and relevant error/timestamp is retained. Hidden reasoning/ciphertext and provider/account identifiers are omitted. Function-call IDs are deterministically pseudonymized to preserve linkage. Continuation input history therefore omits hidden reasoning and uses pseudonymous IDs; the exact initial prompt/settings and all visible content remain unchanged. Do not claim the redacted continuation transport bytes are identical to original private captures. transformation-manifest.json maps original and distributed hashes; public dependent hash references are transparently rebased while immutable frozen originals remain unchanged.

The historical first-run transport retained sanitized error type/code/param but not complete raw error objects. Its correction receipt discloses this limit. New amended attempts explicitly test for generation fields. Capacity failures count as operational availability events, never as model wrong answers or abstentions. Only the first generated answer for each scheduled decision is used. Unknown charge/delivery retains its reservation.

Operational concurrency protocols and any executed segment receipts are retained. The completion phase uses up to 16 simultaneous decisions. Reserved-but-unstarted entries belong to the segment in which they were actually dispatched. Segment-stratified latency is descriptive; differences in concurrency, execution order and source/cache state are not controlled model effects. Inter-phase coordinator pauses and time awaiting local dispatch are excluded from per-decision elapsed time; provider waits inside HTTP requests and bounded retry backoff remain included.

After the eight-worker Flex phase drained, all 81 generated decisions were preserved. The 879 never-generated slots entered a separate availability-completion phase: all 441 remaining Luna requests uniformly use default/Standard service, while 438 remaining Sol requests stay Flex. Prompts, logical IDs, conditions, sources and output limits are otherwise unchanged. Standard has one attempt per turn, with no silent fallback or generated-answer retries. Earlier exhausted capacity histories remain visible. Within-company comparisons separately show same-tier sensitivities and mixed-tier pair counts; tier-equivalent behavior is not assumed.

private/followup-results.json contains the full reviewed summary. The adjacent CSV/JSONL contain all scheduled slots and visible texts. Individual hash-bound review files preserve explanations and observed evidence, including secondary citation/query/source-description concerns that do not automatically change the current-membership score. Operational events with no visible answer text receive an explicit no-answer review; they are not counted as full visible answers. Costs reconstruct frozen Flex and Standard usage/categories/search actions and conservative reservation amounts; no provider billing invoice is claimed. Each grouped cost object covers only its follow-up decisions. Top-level accounting_reconciliation separately adds the original experiment opening amount, reproduces the latest phase ledger, and distinguishes observed conservative charges from pending/unknown-charge reservations. The original experiment contributes its separately reconciled opening amount. Reproducing that original amount from every original request requires the separate v1 complete evidence archive; v2 contains the hash-bound opening receipt and ledger but does not pool the original outcomes with this follow-up. Latency includes the selected execution's tool turns, retry/backoff and local handling. Initial failed attempt durations are retained separately; coordinator pause and queue wait are not model latency.

SanctionsKit designed and funded the study and sells sanctions screening software. This package neither establishes product superiority nor promises that a company is legally safe to transact with.

SanctionsKit does not use AI. Condition T is a custom experimental fixture that lets the tested model query archived official XML; it is not SanctionsKit's interface, API, or a product feature. The operational lesson concerns a defined sanctions-screening process, official-source records and traceability. Employee behavior, employee training outcomes and the production SanctionsKit service were not measured in this experiment.
