API study protocol: current-list answers after official sanctions removals
API v1, 8 October 2026. READY FOR INTERNAL FREEZE, subject to source/gold approval and exact payload commitment. No evaluated API results were used to set this protocol.

Design change and scope

The user requested direct API testing in place of the slower browser execution. The API study is the research project to be completed and published; no separate browser pilot is planned for publication. Preserve all earlier browser captures and their operational deviation internally, but exclude them entirely from API sample sizes, scoring, charts, model comparisons and effect estimates. Disclose in the final method that an earlier browser feasibility phase was discontinued when the user selected API execution. This was a user-requested execution change, not a stopping rule based on whether answers were favorable. The source candidate frame was built before the API study; some entities appeared in earlier browser work. Do not claim that no prior model output on any entity had ever been seen. All API inclusion/ranking decisions and prompts must be locked before evaluated API output.

Question: How do two exactly identified API models answer a specified company's direct membership in the OFAC SDN List at a dated cutoff under three available-evidence conditions? Test companies explicitly removed in three official notices and companies currently named on the same list. Current named-list membership is the outcome, not all sanctions status, ownership-based blocking, permission to transact, business trustworthiness or financial harm. Do not describe absent companies as innocent, exonerated, cleared, safe, operating businesses, or generally unsanctioned. There is no claim about every AI system, all vendors, consumer apps or actual customer screening activity.

Final planned size: 96 distinct corporate entities, consisting of 64 removed and 32 currently listed controls. The removed allocation is 9 from May 28, 27 from July 27 and 28 from October 5, 2026. Controls use fixed program quotas of 20 SDNT, 8 IRAQ2 and 4 SDNTK. Two API models × 96 companies × 3 arms × 2 repetitions = 1,152 scheduled evaluated requests. These are 96 companies, 64 removed companies and three removal events, not 1,152 independent companies. Use exact counts with normal spacing in public copy. No output-driven expansion, case search or subset selection is allowed.

Candidate frame and gold truth

Reuse the original auditable frame of96 eligible organization identities extracted from the SDN-deletion sections of the May28, July27 and October5 notices:9,42 and45 respectively. This is a purposive three-event frame, not a census of all sanctions removals or a random sample of removal events. Later SSI/PLC sections are excluded. July duplicate UID consolidations and surviving-target records are excluded; the ambiguous October identifier-only merged group remains excluded. Preserve the full source frame, all original exclusions and any further identity-review exclusions with reasons before API outputs. Never treat alias paragraphs as separate companies.

API selection seed: sanctionskit-ai-api-study-2026-10-08-v1. Rank eligible identities within each fixed event stratum by SHA-256 of seed, newline, stratum, newline and canonical identity key. Use every eligible May entity, the first27 July entities and first28 October entities. Rank corporate-eligible current source UIDs similarly within the fixed control program strata in SDNT, IRAQ2, SDNTK order, excluding UIDs already selected and any candidate whose normalized name or typed corporate identifier links to another current UID. This general source-only rule was adopted before API outputs after an independently found company/person identifier collision; the original candidate and correction log are preserved. Complete-file lookup packets report actual hit sets, not assumed single-UID matches. Preserve ranked reserves. Control quotas are a design choice, not representative weights or an exact matched design. Selection cannot depend on observed browser or API success/failure. If identity review excludes a selected candidate, record the reason and use the next ranked eligible reserve within that stratum before freeze. If the quota cannot be met, document the reduced panel before any evaluated API request.

Independently downloaded complete official SDN XML, archived official notices, the explicit July deleted/retained-UID table and available prior official snapshots establish the gold source. SanctionsKit's production aggregate is a separate reconciliation subject, never its own gold authority. Archive original bytes, exact URLs, official publication metadata, UTC retrieval times, record count and SHA-256. Local hash integrity is not an official signature unless a matching official digest was actually obtained. Preserve missing historical UIDs and corporate identifiers as unknown.

Verify each identity using every published primary/alias name, stable identifier and any prior UID. Search all current records, not only exact primary-name strings. Review close-name candidates, same-ID collisions, renames, duplicate consolidations and any identified later changes. A genuinely unresolved same-target conflict is excluded before testing. Company eligibility requires corporate evidence beyond a generic Entity tag; current operation is not assumed. The research establishes the published target identities in the inspected records, not exhaustive corporate-successor or ownership relationships. Days since a selected removal notice do not establish continuous absence throughout that interval.

Capture the official current file immediately before the evaluated run and at the end of the complete run, with an additional check if a source-change alert is observed during execution. A six-request dispatch block does not itself trigger a full-source download. If its hash changes, inspect target-level changes. A company with changed membership during the run is temporally contaminated: preserve all outputs, exclude its entire cross-arm/model comparison under this predeclared rule, show the affected counts, and report a sensitivity table. Do not replace it after outcomes exist. Recheck before publication and distinguish later changes from historical evaluation results.

Experimental conditions

For every exact API model, use the same cohort, neutral base question, output schema, reasoning/temperature policy where supported, output-token cap, identity descriptors and number of repetitions. Provider defaults or unavailable settings must be named as such. Each request is fresh, with no previous_response_id, chat history, memory, previous answer, feedback, correction or consumer-app state. Set store:false where supported; do not imply this overrides provider retention terms. Do not expose API credentials in logs, tools, artifacts or public code.

M — Model-only: no retrieval tools and no source packet. The model can answer from its available knowledge or explicitly say CANNOT_VERIFY. A cautious answer is not counted as a false membership assertion. This arm measures behavior under an intentionally constrained information condition; it is not the ordinary behavior of every AI assistant.

W — Live web: same base question, provider web retrieval enabled with a fixed per-request tool-call and output budget. Freeze exact tool names, versions, domain filters (if any), tool-choice configuration and maximum call count. All W requests use the same configuration within the evaluated system. Actual tool traces determine whether retrieval occurred; access enabled alone does not prove use. Do not tune search instructions per company or selectively add the removal notice when a response struggles.

E — Supplied official evidence: same base question plus the frozen identity-specific official-source packet, with web tools disabled. Packets disclose the original official record/notice separately from researcher-derived queries over the complete file. Include exact provenance, timestamps, full query rules and relevant rows; an empty lookup is not an OFAC absence certificate. No gold_label, desired verdict, sponsor success or marketing instruction is supplied.

E deliberately provides decisive dated evidence and removes the need for autonomous retrieval. It is an evidence-comprehension/workflow condition, not proof of production retrieval or of SanctionsKit's screening API. W versus E changes both evidence supply and tool availability; do not call that a clean causal estimate of retrieval alone. M versus W compares available-tool conditions within this fixed task, not every possible real-world workflow. Interpretation must retain those distinctions even if a difference is large.

The two evaluated API request model IDs are gpt-6.1-sol and gpt-6-luna, both from OpenAI. Engineering access checks succeeded before cohort execution. Request IDs are exact API identifiers, not a claim that an alias will remain pinned to one backend indefinitely; preserve returned model metadata. Consumer brand/model equivalence is not inferred. If both models come from one provider, report that fact; do not claim cross-provider coverage. Archive request model IDs and returned model/version metadata, SDK/endpoint, reasoning setting, unsupported settings, tool definitions, temperature/defaults, seed if available, token caps and run times. Do not quietly substitute an unavailable model or promote a display label to an immutable backend version.

Frozen execution settings

Both models use the Responses API, reasoning effort medium, max_output_tokens 4096, service_tier default and store false. No temperature or sampling seed is supplied; unsupported settings and provider defaults remain explicit. M and E expose no tools. W exposes web_search with search_context_size low and external_web_access true, max_tool_calls 4, tool_choice auto, no domain filter, and include web_search_call.action.sources. Actual retrieval usage is read from each returned trace. The same strict JSON schema is used throughout. These are fixed configurations, not each model's claimed optimal configuration. Low search context is a tool setting, not an independently guaranteed small billed-input ceiling.

Prompts, schema and adjudication

The base prompt neutrally asks for the named-list membership of an identified company at the cutoff. It distinguishes historical designation from present membership and allows uncertainty. Identity fields are identical across arms: official name, published location detail and available public corporate identifiers. A source address or registration qualifier is not asserted to be incorporation jurisdiction. The schema requires verdict PRESENT, ABSENT or CANNOT_VERIFY, an explanation, explicit list/date scope and source entries. Source URLs supplied from memory are not treated as retrieved or verified citations. The schema does not force an answer or prohibit an evidence limitation.

Preserve complete raw request, provider response JSON, final text, parsed schema, tool traces, source citations/annotations, response/request IDs, exact UTC start/end, finish reason, token usage, cost receipt/estimate, retry history and hashes. Retain all returned output fields needed to interpret the answer, redacting credentials and account-private data from distributable artifacts. Do not publish hidden provider reasoning; use the actual final answer and allowed tool/citation evidence. Never truncate the stored final answer to fit an attractive excerpt.

The structured verdict is an extraction aid, not unchallengeable truth. Review explanation against it using the frozen rules: a clear final correction controls; conflicting unreconciled current-membership claims are CONTRADICTORY; historical status alone is not a current assertion. CANNOT_VERIFY with no contrary conclusion is UNRESOLVED. An explicit qualified PRESENT/ABSENT remains that membership assertion with assertiveness recorded separately. A correct named-list answer with weak citations stays correct membership, with failed/unverified source support recorded separately. A valid substantive refusal is NO_USABLE_ANSWER; other-list-only discussion is OUT_OF_SCOPE. Scope warnings about ownership or other restrictions do not turn an otherwise clear correct membership answer into an error.

Gold must be reviewed before execution. A separate reviewer labels full responses with model/arm/gold metadata concealed where practicable; record the actual degree of blinding and prior exposure. If the same internal agents already know the cohort, do not claim blind or independent expert review. Preserve both initial labels and adjudication disagreements. No human industry expert review is claimed unless it actually occurs. Provider models may be used to assist parsing, but a model is not the sole unreviewed judge of its own correctness.

Operational errors and hard spending cap

The user authorizes a hard additional-spend maximum of $50 for this API study. Root owns verified pricing, account access and implementation. No request begins until a conservative reservation covers its maximum input, output/reasoning, tool/search and other provider charges plus an explicit safety allowance. Reserve all in-flight requests before dispatch; parallelism must not allow overspend. Count engineering checks, failed/ambiguous attempts and retries against the same $50. Do not rely solely on after-the-fact token totals. If a required model/tool has unbounded or unverified charges under the proposed configuration, do not run it until bounded. Maximum scheduled work is not an instruction to exceed the cap.

Auth/model/schema/tool compatibility checks may use fixed noncohort engineering questions. They must be logged and budgeted, cannot optimize error rates, and are excluded from study outcomes. Freeze tested final configuration before evaluated company requests. If tool support differs between selected models, either choose configurations supporting the declared design before freeze or define and disclose the unequal systems; do not hide a missing web arm.

Generate a deterministic hash-ordered blocked schedule before outcomes using the same fixed seed. Interleave the six source/program strata, with a seeded stratum order and within-stratum company order for each repetition. Complete the scheduled first repetition over all 96 companies before scheduling the second; two repetitions have identical prompts/settings and fresh requests. Each company/repetition block contains all six model-by-arm cells in seeded order and is reserved together before dispatch. Permit at most four active blocks and 24 concurrent requests. Preserve actual dispatch/completion order and acknowledge completion-time variation; the dispatch sequence is fixed rather than selected from answers. A resource stop uses the next complete predetermined block where funds permit, independent of answer content. Report every incomplete cell and actual denominator. No early stop for headline strength, apparent significance or a sponsor win.

There are no automatic retries in the primary schedule. After the entire primary schedule has been attempted, at most one budget-reserved retry may be made for an unambiguous HTTP 429 or 5xx failure if budget remains, in original schedule order. A retry is supplemental recovery and does not erase the original primary failure; show original-attempt analysis and a separately labeled recovered sensitivity analysis. Do not retry a wrong, uncertain, refused, schema-invalid, truncated or otherwise substantive answer. Other transport/provider failures remain primary operational failures; ambiguous delivery is not retried as a new observation. For ambiguous delivery where an answer may already have been generated, use provider idempotency/retrieval if supported or mark operationally unresolved; do not obtain repeated answers until one is convenient. Invalid JSON/length termination is a schema/completion failure, not an automatic model knowledge error. Preserve readable final text for a separately labeled sensitivity review; never silently repair it into a successful structured answer. Duplicate accidental requests are archived/disclosed separately; the first valid scheduled answer remains primary.

Analysis fixed before evaluation

Primary endpoint: wrong PRESENT claims for the64 verified removed identities, reported separately by model and arm with company and response denominators. For each company summarize its two scheduled repetitions; uncertainty is not a wrong-present claim. Report wrong ABSENT claims for the32 current controls as a necessary companion measure. Show complete outcome tables including correct membership, uncertainty, contradiction, out-of-scope, substantive nonresponse and operational failure. Never use an always-absent model's removed-case success as overall accuracy. The artificial64/32 mix is not a real-world prevalence estimate.

Report the fraction of explicit wrong claims over scheduled/observed requests with missingness explicit, plus the number of companies with at least one error in exactly two observed repetitions. Do not compare that latter measure with one-repetition cases without separating them. Show correct-answer coverage and selective correctness together; dropping uncertainty from the denominator alone is misleading. Operationally missing cells get bounds/sensitivity, not silent exclusion that improves a score.

Within each model, prespecified comparisons are M versus W and W versus E; M versus E is secondary. Show paired company-level transitions and paired mean wrong-claim differences. Repeated requests are nested within company; companies are clustered within three purposively selected removal events. Counts and event-stratified case tables are primary. A company-level seeded bootstrap conditional on this finite selected panel may illustrate repeatability uncertainty but cannot supply a population confidence claim and cannot account credibly for only three event clusters. Prefer no inferential p-values in the main report. No causal claim that training memory, web indexing or a particular retrieved page caused an error without separate evidence.

Program, event, country, identity richness, returned source type and specific model contrasts are secondary descriptive analyses. Report the full frozen panel, not only dramatic examples. Do not search for a failure-heavy subgroup as the headline. A zero-error result, small difference, uncertainty-heavy result or sponsor limitation must remain publishable. The title and press release follow measured findings; no first-ever, guaranteed virality, legal clearance, product superiority or industry-wide error-rate promise.

SanctionsKit and public reproducibility

The existing product component is independently evaluated production public-source-data reconciliation: official snapshot hashes, versions, source records and selected identity presence/absence. It is not a hosted screening-API benchmark or an evaluated AI-plus-SanctionsKit integration. Report all mismatches and limitations. Giving a model a researcher-prepared official packet cannot be described as the model using SanctionsKit. A future actual integration/API claim requires its own frozen test and response evidence.

Before publication: freeze manifest and exact request schema/configuration/cohort/evidence/schedule/analysis; run actual API requests; review all results and costs; verify numeric claims and final source stability; publish appropriate public data/code/provenance and redacted final responses; disclose sponsorship, scoped generalizability, evidence-arm asymmetry, engineering exclusions and deviations. Internal freeze is not public preregistration without an actual pre-run public immutable deposit. API findings stand alone; no browser figures are pooled or separately promoted.

Final pre-execution gate: approve the exact 96-case cohort against the refreshed official snapshot, preserve all 64 removed-identity dossiers and 32 control identities, verify bounded all-in reservations, then hash/freeze this protocol, prompts, schema, scoring plan, source manifests, evidence packets, schedule and exact payloads. Evaluation date is 2026-10-08 UTC; record exact source retrieval and every request time. A changed official file before freeze requires source-level and target-level reconciliation, regenerated evidence/metadata where applicable, and renewed gold review; a prior snapshot is never relabeled as current. The internal reviewer has prior cohort/browser exposure and is not a blinded external human expert. Do not read or handle API keys in methods work.
