AI, changing public records, and current sanctions-list membership SanctionsKit research — methodology, version 1 — 8 October 2026 RESEARCH QUESTION Can two API models establish whether a specified business is directly named on the OFAC Specially Designated Nationals and Blocked Persons List when the official record has changed? We compare model-only answers, answers with live web search available, and answers supplied with dated official-source evidence. The outcome is membership in this one named list on the evaluation date. It is not a finding about every applicable sanction, ownership-based blocking, permission to transact, or a business's trustworthiness or present operations. PANEL AND SOURCE OF TRUTH The panel contains 96 business identities: 64 removed identities from three selected OFAC notices and 32 currently named controls. There are nine removals from the 28 May 2026 notice, 27 from 27 July, and 28 from 5 October. Control program quotas are 20 SDNT, eight IRAQ2 and four SDNTK. These quotas are design choices, not population weights. The removal events were purposively selected; this is neither a census of removals nor a representative sample of all sanctions-screening questions. The eligible removal frame contained nine, 42 and 45 identities respectively. Selection within each stratum used SHA-256 ranking with the fixed seed sanctionskit-ai-api-study-2026-10-08-v1. The published cohort and selection code make the selected identities inspectable. Aliases were grouped with their target identity. Duplicate-record consolidation was not treated as removal of the target. Six corporate targets in a July duplicate-UID consolidation were excluded, as was an unresolved October merged-identifier group. Before any evaluated API answer, one initially selected control was excluded because its corporate identifier also matched an individual record; the next deterministically ranked eligible control replaced it. No API outcome influenced these decisions. Gold labels were established from independently downloaded complete official SDN XML, the original OFAC notices, published aliases and corporate identifiers, and an available prior snapshot where applicable. Known names and typed identifiers were queried across every current record, including aliases. Close-name candidates and identity collisions were reviewed. A second XML parsing and linkage check reproduced the expected status for all 96 selected identities. These are public-record identity determinations, not exhaustive investigations of corporate successors or ownership. The API source snapshot was retrieved on 8 October 2026 at 17:08:39.800127 UTC. Its publication date is 8 October 2026, it contains 19,416 records, and its SHA-256 is 259336aa866dda3fc0cd3cddb933f4b5d48025d47035bcdece9b4ec90bf7ed92. The source URL is https://sanctionslistservice.ofac.treas.gov/api/PublicationPreview/exports/SDN.XML. The official file changed during preparation; all packets were rebuilt against the new file before the API study was frozen. The selected identities and their membership labels did not change. EXPERIMENTAL DESIGN Models: gpt-6.1-sol and gpt-6-luna, through the OpenAI Responses API. These are two API models from one provider. Results do not describe every AI system, other providers, or the consumer ChatGPT interface. Requested and returned model IDs are retained. An API model name is not a guarantee that the provider will keep its underlying implementation unchanged indefinitely. Every company receives the same base question under three conditions, with two fresh repetitions per model and condition. This schedules 96 × 2 × 3 × 2 = 1,152 requests. It remains a panel of 96 companies, with 64 removals clustered in three events; repeated answers are not independent companies. M — Model-only. No external tools and no source packet. The model may answer or explicitly state CANNOT_VERIFY. This intentionally constrained condition does not represent all ordinary assistant use. W — Live web available. The same question, with the provider's web_search tool, search_context_size low, external_web_access true, automatic tool choice, and max_tool_calls 4. No domain filters are supplied. The response retains actual tool actions and source URLs. Tool availability is distinguished from observed tool use, and a URL appearing in a trace is not proof that its content supports the answer. Requested limits and actual returned traces are both retained. E — Supplied official evidence. The same question plus an identity-specific source packet; external tools are disabled. Packets contain official provenance, relevant record or removal-notice excerpts, and a clearly labeled researcher-derived lookup over the complete archived XML. Empty query results are not described as an OFAC-issued absence certificate. No desired verdict or product claim is supplied. This condition gives the model decisive evidence and tests its interpretation; it does not test autonomous retrieval or an AI integration using the SanctionsKit API. W versus E changes both the supplied evidence and tool availability, so their difference is not an isolated causal effect of retrieval alone. Common settings: reasoning effort medium; max_output_tokens 4,096, including reasoning; store false; default service tier; no conversation history or previous_response_id. Temperature and sampling seed were not overridden. Returned provider settings are retained. The as-of date appears in the question. The prompt requires a bounded list-membership answer, permits uncertainty, distinguishes historical designation from current membership, and prohibits claiming retrieval that did not occur. These instructions are deliberately careful; results are conditional on them, rather than on an unspecified casual prompt. The full schedule, exact prompts, schema, source packets, selection and analysis plan were internally frozen at 17:17:52.794650 UTC, before the first evaluated API request. The commitment SHA-256 is b1a3c86878e70bd7d6bb7fdf3aa3bce44826b5ab059ab734d281073df28f104c. This was an internal pre-execution commitment, not a public preregistration. Company blocks balance the six model-condition cells, with deterministic stratum interleaving. All first repetitions precede second repetitions in the dispatch schedule. Maximum concurrency is 24 requests in four six-cell blocks. An earlier browser feasibility phase was discontinued when API execution was selected. Some API-panel identities had appeared in that phase; the panel was not wholly unseen before API design. Its responses are excluded from the study, all denominators and all comparisons. Two non-company API engineering checks verified access and response capture; they are also excluded. No API-result-dependent prompt tuning, replacement, additional company search, or early stopping for an attractive finding was allowed. SCORING AND REVIEW The strict response schema requires PRESENT, ABSENT or CANNOT_VERIFY, a dated explanation, sources and limitations. Deterministic code compares the structured verdict with the frozen gold. Full visible final answers are then inspected for disagreement between verdict and explanation, unresolved contradiction, inappropriate scope and source-access claims. Review is internal to this vendor-produced study; it is not blind external peer review or human expert certification. Primary outcomes are wrong PRESENT claims about removed identities and, separately, wrong ABSENT claims about current controls. Correct membership, uncertainty, contradiction, out-of-scope replies, substantive nonanswers and operational failures remain distinct. A correct status with weak citations stays correct for the membership outcome and is flagged separately for source support. Not every returned URL and its content was independently fetched; citation flags are not a complete citation-support benchmark. Consequential wrong claims receive additional source corroboration. A historical designation alone is not a current-membership claim. A generic ownership warning does not turn a correct named-list answer into an error. The first scheduled attempt is primary. Transport failure, rate limit, invalid schema and length termination are operational outcomes, not automatically wrong membership answers. Any clearly failed HTTP 429/5xx retry would be supplemental, retain the original failure and never replace a substantive answer. Exact prompts, visible answers, tool metadata, timestamps and scoring records are retained. Provider reasoning ciphertext is not decoded or published. Counts are reported by model, condition and source cohort, with event-level tables and paired company summaries. Company error counts require the two scheduled repetitions to be observed. No population accuracy estimate, independent-observation p-value or confidence interval is claimed. The artificial 64/32 case mix must not be interpreted as real-world prevalence. Correct-answer coverage is reported alongside wrong answers and uncertainty, rather than dropping uncertainty to produce a flattering score. SOURCE STABILITY AND PRODUCT COMPARISON The full official file was checked before testing and after the complete API run. The end check completed at 17:48:43.657357 UTC on 8 October 2026 and returned bytes identical to the frozen snapshot; its SHA-256 remained 259336aa866dda3fc0cd3cddb933f4b5d48025d47035bcdece9b4ec90bf7ed92. The public stability receipt records that check. A selected identity whose true membership changed during the run would have had its entire cross-condition comparison excluded under the predeclared temporal-contamination rule, with every output retained and affected counts disclosed. No such exclusion was required. Later publication checks are separate from the historical evaluation cutoff. SanctionsKit's product component is a separate read-only reconciliation of its active production public-source records. It does not query customer screening activity and is not a hosted screening-API accuracy benchmark, fuzzy-matching test, or tested AI-plus-SanctionsKit integration. Selected-identity agreement and full inventory/version agreement are reported separately, including any mismatch. The documented OFAC refresh schedule is six hours, with a twelve-hour freshness cutoff since a successful check; it is not a real-time update promise. Neither matching this panel nor accepting a supplied source packet establishes universal product accuracy. At the initial preparation observation, 17:10:51.995 UTC, production still exposed the 5 October source: the selected 96 identities agreed, but the full inventory lacked 55 current records and retained two removed vessel records. Its most recent successful check was 12:36:22.384 UTC, within the documented refresh interval. At the end observation, 17:48:44.634 UTC, production recorded source retrieval at 17:11:22.484 UTC and matched the 8 October official source hash and all 19,416 source UIDs. All 96 selected identities agreed in membership, name/identifier lookup results and the compared common record fields. The researchers did not trigger a production refresh. These observations establish the measured state at those times, not uninterrupted parity between observations. REPRODUCTION, LIMITATIONS AND CORRECTIONS The public package contains the selected company facts, official notice links, source fingerprints, exact prompts and configuration, all visible first-attempt responses, tool-source metadata, scoring decisions, aggregate tables and reproduction code. A transformation manifest maps private raw-capture hashes to sanitized public captures. Provider/account identifiers and reasoning ciphertext are omitted; final answer text and scientific request content are preserved. No customer data or credentials are included. The package reproduces the published scoring and exposes selected source evidence. Complete original government downloads are retained in the research archive; the live government download can change. For access to the archived original source files, contact support@sanctionskit.com and specify the study and source hash. A file hash identifies the archived bytes but is not an official digital signature and does not by itself make a later live download identical. The exact historical full-file absence audit requires the archived original; selected excerpts alone do not independently prove complete-file absence. This availability limit is explicit rather than presented as unrestricted end-to-end replication. SanctionsKit, a sanctions-screening software provider, produced and funded the study. That commercial interest is relevant to interpretation. The source panel, model/provider scope, careful prompts, limited tool budget, evidence-arm advantage, repeated-company structure and internal review all limit generalization. No measured commercial harm, legal clearance, cross-provider superiority or universal model failure is asserted. Figures and selected findings may be cited with attribution to SanctionsKit and a link to the full study; identified corrections should be sent to support@sanctionskit.com. Substantive corrections will be versioned and explained rather than silently replacing the evidence. OFFICIAL REFERENCES OFAC 28 May notice: https://ofac.treasury.gov/recent-actions/20260528 OFAC 27 July notice: https://ofac.treasury.gov/recent-actions/20260727 OFAC 5 October notice: https://ofac.treasury.gov/recent-actions/20261005 OFAC ownership guidance, FAQ 401: https://ofac.treasury.gov/faqs/401 OpenAI model documentation: https://developers.openai.com/api/docs/models/gpt-6.1-sol OpenAI model documentation: https://developers.openai.com/api/docs/models/gpt-6-luna OpenAI web-search documentation: https://developers.openai.com/api/docs/guides/tools-web-search