SanctionsKit research · October 2026
When AI gets sanctions status wrong: a 96-company test
A 96-company study checked Google AI Overviews and two separate API experiments against official OFAC records. The findings show why sanctions screening needs current evidence and a defined review process.
On this page
Google’s AI Overview said no listing. The official file contained the company.
A manual company check can begin with a web search and put an AI-generated answer in front of the reviewer. SanctionsKit’s position is that the underlying screening checks need deterministic, reproducible rules: record the source version, apply the approved matching policy and keep the evidence connected to human review.
We searched Google for the same 96 company identities and saved the actual AI Overview text, source cards and screenshots. This consumer-search observation is separate from both API experiments below.
For NEW AHMADI LTD. in Afghanistan, the Overview said there was no public SDN listing. The complete official file contained that exact name as entity 13128, with two aliases and an Afghanistan address. General advice to double-check details did not retract the answer’s false absence claim.
Across the 96 searches, 68 Overviews gave a correct current-membership answer and 17 gave a wrong one: 14 false listing claims and 3 false absence claims. The remaining outcomes are shown separately below. These are observations from a deliberately selected panel, not an error rate for Google searches generally.
What the Google AI Overview said
Google Search AI Overview · 8 October 2026
Based on official records from the OFAC Sanctions List Search database, there is no public listing for an entity named "NEW AHMADI LTD." or "New Ahmadi Ltd." registered in Afghanistan on the Specially Designated Nationals and Blocked Persons (SDN) List.Read the complete captured answer and evidence
What the official record contained
Complete OFAC SDN XML · 8 October 2026
UID: 13128 Name: NEW AHMADI LTD. Aliases: NEW AHMADI COMPANY LTD; NEW AHMADY LTD Program: SDNTK Address: Sarafi Market, Shop 48/49, Gereshk, Helmand, AfghanistanInspect the saved official record and identity check
Read the complete question, answer context and official identity check
The exact query was “Is NEW AHMADI LTD. (Afghanistan) on the OFAC SDN List?” The answer then recommended checking exact registration details, addresses and aliases, and offered help if more details were provided. It did not say that its opening no-listing claim was uncertain or withdraw it.
The official record shown above is a labeled transcription of fields from the archived XML, not a prose quotation from OFAC. The primary name and Afghanistan location match the query; both aliases occur in the same record. We inspected entity 13128 directly and searched the complete 19,416-record file. The file contained no corporate registration number for this entity, so none is supplied here.
The source link displayed in the answer used an opaque Google redirect. The captured link and surrounding organic results are preserved, but they do not establish the exact page body the answer used. The finding is the observable false absence assertion, not a claim about the system’s internal reasoning.
The reporter examples preserve the complete generated answer, source context, exact query and official identity evidence. The Google-only evidence archive includes all 96 observations and the complete reference file. This illustrative case was selected after review, not as a random or typical search.
Give every employee a defined screening process
SanctionsKit does not use AI. Companies can establish their screening process through the SanctionsKit interface or integrate the sanctions screening API into their own application. Each result records the selected coverage, source versions and matching rules. Retained evidence connects the submitted party, screening outcome and human review in one traceable record.
Train employees to follow that defined flow. An unresolved or unanswered web check could be mistaken for “nothing found” if the workflow leaves its meaning to the reader. That is a possible process risk; these experiments did not measure employee behavior. The false absence claim above is a different observed problem: a definite answer that contradicted the current official record.
SanctionsKit keeps potential matches, completed no-match outcomes and incomplete checks distinct. Its Review queue makes the work visible as “Needs review”, “Awaiting second review” or “Review complete”, with assignment and recorded decisions. These controls give teams an explicit status, owner and next step for managing screening risk.
OFAC’s Framework for Compliance Commitments calls for clear procedures, escalation, recordkeeping and role-appropriate training. Our recommended screening process applies those principles to the source-checking problem illustrated here.
- Define where checks happen: use the SanctionsKit interface or the approved API integration with a published screening policy. Retain the source coverage, data versions and matching policy used for each check.
- Choose matching sensitivity deliberately: Standard and Focused policy profiles let the team balance fuzzy-name alert volume against missed name variants. Standard includes more fuzzy-name alerts; Focused requires closer similarity. Evaluate both on representative cases, including unnecessary flags and missed candidates.
- Work the Review queue: standard-retention potential matches automatically create review work. Assign the investigation, compare company identifiers and source facts, record the reasoning, and complete required independent approval. Policy-required wider checks can also send a no-match result into review.
- Retain the evidence: preserve submitted details, source versions, screening results and review decisions under your retention policy. The screening audit trail guide explains how to retrieve that path.
- Train employees to follow the flow: show them where to submit a check, how to distinguish a completed no-match from an incomplete check, who resolves an uncertain identity and how to record the decision. Rehearse the reviewer handoff before using it for real checks.
A removal notice also appeared beside a false listing claim
The error ran in the other direction for KEENCLOUD LIMITED. Its Overview said the company was listed and displayed OFAC’s 27 July removal notice among its sources. The exact London company was in that notice’s deletion section and absent from the complete current file.
What the Google AI Overview said
Google Search AI Overview · 8 October 2026
Yes, KEENCLOUD LIMITED (located at 11 Catherine Place, Westminster, London, United Kingdom) is listed on the U.S. Department of the Treasury's Office of Foreign Assets Control (OFAC) Specially Designated Nationals and Blocked Persons (SDN) List under the Iraq sanctions program ([IRAQ2]).Read the full answer and displayed source cards
What the official notice said
OFAC notice · 27 July 2026
The following deletions have been made to OFAC's SDN List: KEENCLOUD LIMITED, 11 Catherine Place, Westminister, London, United Kingdom [IRAQ2].Read the official OFAC removal notice
Read the KEENCLOUD context and source boundary
The exact query was “Is KEENCLOUD LIMITED (United Kingdom) on the OFAC SDN List?” The answer repeated the company name, address and IRAQ2 program, then pointed to OFAC and Federal Register notices. It did not qualify its present-listing claim as historical or unresolved.
The official quotation joins the deletion heading with the complete company entry; other deletion entries intervene. The address is spelled “Westminister” in the official record and “Westminster” in the Overview. The name, street and country identify the same company. An independent full-file name and alias check found no current entry.
The screenshot and DOM preserve a source card titled “Sanctions List Removals; Sanctions List Updates” dated 27 July, plus a later Federal Register notice. We opened the preserved OFAC citation and confirmed that it resolved to the official 27 July removal page. That confirms the citation destination, but does not reveal the exact passage delivered to the generator. We report the contradiction without claiming a proven internal cause.
Every Google search outcome, including history and no Overview
Only the generated Overview prose was scored. Source-card snippets and ordinary search results were kept separate. A historical statement can be accurate without explicitly answering current membership: “was removed” was retained as historical-only, not treated as wrong. An abstention leaves current status unresolved; a missing Overview is an availability outcome.
- Correct current answer
- Wrong current answer
- Abstention
- Historical only
- Other scope only
- No Overview
All selected companies
Removed companies
Listed controls
Read the Google collection method, timing sensitivity and limits
We fixed 96 exact company-name/location queries and their order before collection. Each asked whether that company was on the OFAC SDN List. Researchers recorded the actual Google Search interface, expanded available Overview text and captured displayed citation context. This is an observation of a consumer product, not an API model test or a simulation of Google answers.
Collection used four tabs in fixed batches. The archive discloses an access interruption, viewport differences and eight observations without recorded submission timestamps, four of which also lack first-results timestamps. Missing timestamps were not estimated. The one no-Overview result was confirmed 37 seconds after first results, beyond the planned 30-second observation window; it was not scored as a wrong or unanswered generated claim.
The complete-timing subset contains 88 searches: 64 correct current answers, 14 wrong current answers, one abstention, seven historical-only answers, one broader-scope answer and one search without an Overview. We report all 96 searches and distinguish the 95 displayed Overviews.
The complete official SDN XML was unchanged in source checks before and after Google collection. Each full generated answer received an internal, unblinded review; every wrong claim received a separate internal source and identity check. Ambiguous historical and broader-list wording was also adjudicated. This was not external peer review.
The original API study and evidence were already public before Google collection. SanctionsKit record pages appeared in some organic results, but no reviewed visible AI citation pointed to our study. We did not resolve every opaque citation destination, so indirect influence cannot be ruled out.
Google answers are dynamic and may differ by time, location, interface and search context. These are preserved observations for a selected panel, not a population sample, a consumer comparison with the API models, or evidence of employee behavior. Direct membership agreement does not establish adequate source evidence, ownership screening or transaction permission.
Download the separate Google archive, manifest and reviewed results. The included offline verifier checks source versions, captured text and image bindings, reviewed spans and counts.
Scroll horizontally to compare every column.
| Panel | Correct | Wrong | Abstention | Historical only | Other scope only | No Overview |
|---|---|---|---|---|---|---|
| All selected companies | 68 | 17 | 2 | 7 | 1 | 1 |
| Removed companies | 42 | 14 | 1 | 6 | 0 | 1 |
| Listed controls | 26 | 3 | 1 | 1 | 1 | 0 |
The original API experiment: a removal notice became a false listing claim
OFAC’s 5 October notice put COMERCIAL JINAN S.A. under SDN deletions. Asked about its status on 8 October, gpt-6-luna twice cited that notice to say the company was still listed. The name and Guatemala tax identifier matched. A check of the complete current file confirmed that the company was absent.
The models differed sharply. With web search, gpt-6.1-sol gave 78 correct answers, one wrong answer and 113 that could not verify. gpt-6-luna gave 43 correct answers, 28 wrong answers and 121 that could not verify. Each model answered 192 questions: the same 96 companies, twice.
These figures describe the completed 1,152-response experiment. A separate, reviewed 960-decision follow-up below tests more web capacity, original documents and model-chosen queries against the complete official file. Its results are not pooled with the first experiment.
What the official notice said
OFAC notice · 5 October 2026
The following deletions have been made to OFAC's SDN List: COMERCIAL JINAN S.A., 20 Calle No. 16-36, Proyecto 4-4, Zona 6, Guatemala, Guatemala; NIT # 4151952-3 (Guatemala) [SDNT].Read the official OFAC notice
What the AI answer said
gpt-6-luna · live web search · 8 October 2026 · second attempt
PRESENT. OFAC's October 5, 2026 SDN List update names COMERCIAL JINAN S.A. with the matching Guatemala NIT # 4151952-3 and Guatemala location. This supports direct membership in the named SDN List as of October 8, 2026.Read the saved answer and source links
Read the full example context and all 12 outcomes
The model question was direct membership in the OFAC SDN List as of 8 October 2026, using the company’s name, Guatemala location and NIT 4151952-3. The notice heading above is followed by other deletion entries; the quoted company entry is reproduced in full. The excerpt of the model answer preserves its complete explanation before the source citation. Its structured verdict was PRESENT.
The answer linked the 5 October removal notice as current-membership evidence and a 9 July 2009 designation notice as historical evidence. Its limitation said that the answer concerned direct named-list membership, rather than ownership restrictions, other lists or transaction permission. Those caveats did not correct the false membership conclusion.
The reference check searched known names, aliases and the supplied corporate identifier across the complete official file published on 8 October. It found no matching current record. Both models correctly answered ABSENT in both repetitions when given the prepared official-evidence packet and current-file lookup.
The complete example record includes all response text, limitations and citations for both models, all three conditions and both repetitions. It also records the official source check. This case was selected to illustrate a false listing claim that contradicted its cited removal notice, not as a random or typical case.
Scroll horizontally to compare every column.
| Model and test condition | First answer | Second answer |
|---|---|---|
| gpt-6.1-sol · Model only | Cannot verify | Cannot verify |
| gpt-6.1-sol · Live web search | Cannot verify | Not listed (correct) |
| gpt-6.1-sol · Supplied official evidence | Not listed (correct) | Not listed (correct) |
| gpt-6-luna · Model only | Cannot verify | Cannot verify |
| gpt-6-luna · Live web search | Listed (wrong) | Listed (wrong) |
| gpt-6-luna · Supplied official evidence | Not listed (correct) | Not listed (correct) |
The models behaved differently
Luna produced 28 of the 29 wrong web-search answers. All 29 claimed that a removed company was still listed; they concerned 23 companies, with some receiving the same wrong answer twice. Neither model falsely called a listed control absent. These counts describe a selected test panel, not the frequency of mistakes across all AI use.
Each bar below includes all 192 answers for one model and condition. The 96 companies comprise 64 removed identities and 32 listed controls, each asked twice. The official reference status was known for every company; “cannot verify” records the model’s decision to withhold a definite answer.
The supplied-evidence condition produced 382 correct answers and two unresolved answers. Its packets included official excerpts and our completed lookup against the current file. That tests the model’s use of prepared evidence; retrieval and lookup were performed before the model answered.
- Correct
- Wrong
- Cannot verify
gpt-6.1-sol · Model only
gpt-6.1-sol · Live web search
gpt-6.1-sol · Supplied official evidence
gpt-6-luna · Model only
gpt-6-luna · Live web search
gpt-6-luna · Supplied official evidence
See every result split by removed companies and listed controls
A wrong answer for a removed company is a false PRESENT claim; a wrong answer for a listed control would be a false ABSENT claim. Correct answers, wrong claims and uncertainty are separate. All 1,152 scheduled requests completed, with no API failures, missing responses or retries.
The original cohort-split chart, company-level repeat chart and aggregate results preserve the complete counts.
Scroll horizontally to compare every column.
| Model and evidence | Removed companies (128 answers) | Listed controls (64 answers) |
|---|---|---|
| gpt-6.1-sol · Model only | 0 correct · 0 wrong · 128 cannot verify | 0 correct · 0 wrong · 64 cannot verify |
| gpt-6.1-sol · Live web search | 27 correct · 1 wrong · 100 cannot verify | 51 correct · 0 wrong · 13 cannot verify |
| gpt-6.1-sol · Supplied official evidence | 128 correct · 0 wrong · 0 cannot verify | 64 correct · 0 wrong · 0 cannot verify |
| gpt-6-luna · Model only | 0 correct · 0 wrong · 128 cannot verify | 0 correct · 0 wrong · 64 cannot verify |
| gpt-6-luna · Live web search | 11 correct · 28 wrong · 89 cannot verify | 32 correct · 0 wrong · 32 cannot verify |
| gpt-6-luna · Supplied official evidence | 126 correct · 0 wrong · 2 cannot verify | 64 correct · 0 wrong · 0 cannot verify |
Answering more and answering correctly are different measures
An assistant can withhold a decision or commit to a status. These two measures keep that distinction visible: how often the model gave a definite answer, and how often that definite answer was correct. The second percentage excludes uncertainty by definition and should always be read beside the first.
Scroll horizontally to compare every column.
| Model | Gave a definite answer | Correct among definite answers |
|---|---|---|
| gpt-6.1-sol | 79 / 192 (41.1%) | 78 / 79 (98.7%) |
| gpt-6-luna | 71 / 192 (37.0%) | 43 / 71 (60.6%) |
What we asked and what each model could access
The original experiment compared model-only answers, bounded live web search and supplied evidence across 96 companies, two models and two repetitions.
Read the original question, instructions and evidence conditions
The question began: “As of 2026-10-08, is this specific company named on the U.S. Treasury OFAC Specially Designated Nationals and Blocked Persons List (SDN List)?” We supplied its official name and available identity details. The instructions said: “If identity or current membership cannot be established, answer CANNOT_VERIFY.”
We tested gpt-6.1-sol and gpt-6-luna with medium reasoning effort and a 4,096-token output limit. Each company was tested twice in each condition, using a fresh conversation every time. That gives 96 companies × 2 models × 3 conditions × 2 repetitions = 1,152 responses.
Model only: no browsing tools or current-source material were provided.
Live web search: browsing was available, with a maximum of four tool calls and the provider’s low search-context setting.
Supplied official evidence: browsing was disabled. We provided relevant official excerpts and our completed lookup against the complete current list, including whether a matching record was found. This condition tests the interpretation of that packet.
What “cannot verify” tells us
The models were told to avoid guessing about current membership. Without tools or supplied evidence, all 384 responses declined to verify. That was an expected consequence of asking for current evidence without providing a way to obtain it.
With web search available, 234 of 384 answers could not verify. Some described finding an old designation or removal notice but being unable to inspect current-list data. Of those 234 responses, 222 completed all four permitted web actions. That overlap does not establish that the tool cap caused the abstention or that a larger budget would resolve it.
The two unresolved answers with supplied evidence raised concerns about missing company identifiers. Unresolved results were explicit model decisions, not failed API requests. The original experiment did not test unrestricted browsing.
What changed when the models could query the actual list?
In a separate follow-up, the models chose their own name and identifier queries against the complete archived official file. Sol gave 95 correct answers and one “cannot verify”; Luna gave 91 correct answers and five “cannot verify”. Neither made a wrong membership claim in that condition. In the fresh four-action web baseline, Sol gave 39 correct answers and 57 “cannot verify”; Luna gave 22 correct, eight wrong and 66 “cannot verify”. Each count is out of 96 companies.
This comparison makes the source-checking problem concrete: a complete, dated list can support a reproducible check in a way that a name appearing in a web answer cannot. The model-operated lookup was a research fixture. The company process recommended above uses SanctionsKit’s interface or API, official records and a documented path to human review.
The follow-up scheduled 960 decisions: the same 96 companies, two models, five conditions and one repetition. It is separate from the original 1,152-response experiment. Processing tiers were mixed for Luna; the details preserve those differences and every scheduled outcome.
- Correct
- Wrong
- Cannot verify
- Truncated
gpt-6.1-sol · Web: four actions, low context
gpt-6.1-sol · Web: eight actions, low context
gpt-6.1-sol · Web: four actions, high context
gpt-6.1-sol · Official documents supplied
gpt-6.1-sol · Archived-list lookup tool
gpt-6-luna · Web: four actions, low context
gpt-6-luna · Web: eight actions, low context
gpt-6-luna · Web: four actions, high context
gpt-6-luna · Official documents supplied
gpt-6-luna · Archived-list lookup tool
Read follow-up methods, evidence limits and answer rates
The three web conditions used the same current-membership question: four actions with low search context, eight actions with low context, and four actions with high context. These are maximum tool-call settings, not a requirement to use every action. All conditions used medium reasoning effort and a total 4,096-token output allowance. A new conversation was used for each decision.
The documents-only condition supplied a removed company’s original removal notice, or a listed control’s current XML record. It did not include our completed lookup. A historical removal does not alone prove absence on a later date, so many abstentions in this condition were appropriate. A correct membership label is also distinct from adequate current-source support: the review records cases whose conclusion matched the reference but whose evidence was only historical.
The lookup condition let each model query the complete 19,416-record archived official XML by normalized exact names and aliases, or typed corporate identifiers. It did not supply the reference answer. It allowed at most four function calls and preserved each query and result. The tool’s free-text identifier-type field did not enumerate accepted values: some country-qualified labels were rejected even though plain NIT was supported. Those interface limitations and subsequent fallback queries are recorded separately from membership correctness.
A deterministic lookup of the panel agreed with all 96 reference labels. This checks reproducibility on the selected panel; it is not an independent accuracy estimate because selection and reference checking used related lookup procedures. The experimental tool was not the SanctionsKit product.
Testing ran on 8 October 2026, from approximately 20:15 to 21:42 UTC. Sol used Flex for all 480 decisions. Luna used Flex for 39 decisions and Standard/default for the remaining 441, following an availability amendment applied only to never-generated slots. All 81 answers generated before that amendment were preserved. No generated answer was retried or replaced. Source downloads before and after the follow-up matched the pinned official XML bytes.
All 960 slots are retained: 958 completed outputs and two truncated outputs. Reviewers read all 959 outputs containing visible text, including the partial answer; the remaining no-text event received an operational review. All 32 wrong membership claims received a second source check. Review was internal and unblinded. The 309 documented capacity rejections were operational events, not model answers, mistakes or abstentions.
The follow-up’s usage-based cost estimate was $17.74. Conservative combined accounting for the original experiment and follow-up was $40.72, with no unsettled reservations; these are recorded-usage estimates, not a provider invoice. The archive separates processing phases and latency definitions. Queue pauses, concurrency changes, capacity backoff and mixed tiers prevent a controlled speed comparison.
The follow-up results, separate evidence archive and manifest preserve every condition, cohort, request, visible answer, tool result, source check and review. These are descriptive results for a selected panel, not an estimate of industry-wide error rates.
Scroll horizontally to compare every column.
| Model and condition | Gave a definite answer | Correct among definite answers |
|---|---|---|
| gpt-6.1-sol · Web: four actions, low context | 39 / 96 (40.6%) | 39 / 39 (100.0%) |
| gpt-6.1-sol · Web: eight actions, low context | 55 / 96 (57.3%) | 55 / 55 (100.0%) |
| gpt-6.1-sol · Web: four actions, high context | 46 / 96 (47.9%) | 46 / 46 (100.0%) |
| gpt-6.1-sol · Official documents supplied | 65 / 96 (67.7%) | 65 / 65 (100.0%) |
| gpt-6.1-sol · Archived-list lookup tool | 95 / 96 (99.0%) | 95 / 95 (100.0%) |
| gpt-6-luna · Web: four actions, low context | 30 / 96 (31.3%) | 22 / 30 (73.3%) |
| gpt-6-luna · Web: eight actions, low context | 45 / 96 (46.9%) | 35 / 45 (77.8%) |
| gpt-6-luna · Web: four actions, high context | 35 / 96 (36.5%) | 21 / 35 (60.0%) |
| gpt-6-luna · Official documents supplied | 36 / 96 (37.5%) | 36 / 36 (100.0%) |
| gpt-6-luna · Archived-list lookup tool | 91 / 96 (94.8%) | 91 / 91 (100.0%) |
More web searching did not remove the need to check the record
Under the higher web-action allowance, some previously unresolved questions received correct answers. Sol’s correct answers rose from 39 to 55 when the action limit increased from four to eight, with no wrong claims in either condition. Luna’s correct answers rose from 22 to 35, while wrong claims rose from eight to ten and two outputs were truncated. The larger-context condition produced 46 correct answers for Sol and 21 for Luna; Luna made 14 wrong claims in that condition.
A recent publication date is not necessarily a recent sanctions action. In the larger-context condition, Luna said ATLAS EQUIPMENT COMPANY LIMITED was listed in an official notice published on 7 October. The cited notice actually records its removal on 27 July. A citation and a fresh date still need to be checked against the entry, its heading and the effective action.
Compare the same companies and inspect the ATLAS example
For Sol, all pairs used Flex. Moving from four to eight web actions changed 18 companies from cannot verify to correct and two from correct to cannot verify. For Luna, among the 85 complete Standard/default pairs, 15 changed from cannot verify to correct and seven from cannot verify to wrong. Four additional eligible pairs used different tiers, five used Flex in both conditions, and two were excluded for truncation. These are paired descriptive observations from one repetition, not proof of a general causal effect.
With larger context, Sol changed 11 companies from cannot verify to correct and four in the opposite direction. Among Luna’s 87 Standard/default pairs, nine changed from cannot verify to correct and nine from cannot verify to wrong. Three further pairs used different tiers and six used Flex. The full transition matrices, including wrong-to-correct and correct-to-uncertain changes, are retained in the results download.
Luna’s full ATLAS explanation was: “PRESENT. OFAC’s official notice published October 7, 2026 lists ATLAS EQUIPMENT COMPANY LIMITED, with a London, United Kingdom address, among entries on the SDN List. The exact company name and UK location match the supplied identity details.” Its scope limitation concerned direct membership, other lists, ownership and transactions; it did not qualify that PRESENT claim.
Section A of the cited Federal Register notice states that the companies were removed on 27 July. It includes “ATLAS EQUIPMENT COMPANY LIMITED, 55 Roebuck House, Palace Street, London, United Kingdom [IRAQ2].” That entry appears before the separate section B on duplicate removals. The original OFAC removal notice and the complete current XML check corroborate absence.
This is a second selected illustration of an answer contradicting its cited source. It demonstrates the difference between publication date and action date. The experiment retained citations and web-action traces, but not the retrieved web-page bodies; it does not establish exactly which passage the model received or the internal cause of the error. The archive includes the full answer and independently obtained source corroboration.
How the companies and answers were checked
The panel contains 64 deliberately selected removed companies and 32 listed controls. Official source records, published selection rules and retained responses make each result inspectable; the panel does not represent the mix in a normal screening queue.
Read company selection, original execution and study limits
The removed group contains 64 companies from three OFAC announcements: 9 from 28 May, 27 from 27 July and 28 from 5 October 2026. We selected these announcements deliberately, then used a fixed SHA-256 ranking within each eligible group before the API test. Alias deletions and duplicate-record merges were not treated as removal of a company; unresolved identity matches were excluded.
The 32 listed controls were selected using fixed program quotas: 20 SDNT, 8 IRAQ2 and 4 SDNTK. That mix provides checks in both directions; it is not the mix expected in a normal screening queue. Two answers about a company are still one company, and multiple companies share each removal notice. These are descriptive counts for a selected panel, not an industry-wide error rate.
Luna gave at least one wrong web-search answer for 23 companies: five received it twice and 18 received it once, producing 28 wrong answers. Sol’s one wrong answer concerned a company already in that group of 23. The worked example shows a listing claim that contradicted its cited removal notice; other wrong answers relied on historical listing pages. We have not attributed all 29 errors to the same cause.
The reference was the complete official SDN file published on 8 October, with 19,416 records, retrieved at 17:08:39 UTC. Names, aliases and available corporate identifiers were checked, with a second parsing and identity check. The source changed during preparation, so packets were rebuilt before testing. A fresh download after testing returned identical bytes.
The questions, company selection, source packets and scoring rules were fixed before the evaluated API requests. Testing ran from 17:18 to 17:47 UTC on 8 October. All 1,152 requests completed; there were no missing responses, API failures or retries. Each final answer was checked against the reference status, and each wrong membership claim received an additional source check. This was an internal review, not external peer review.
An earlier browser exercise was abandoned and excluded from these results. Some of the same company names had appeared in that preparation work. The API study was therefore not conducted on an entirely unseen set of companies. No evaluated API answer was replaced or omitted to improve the results.
The original API experiment was a dated test of direct named-list membership, including a condition with the provider’s web-search tool. The separate Google observation above tested consumer search output. A missing SDN entry does not establish that a transaction is permitted; ownership rules, other lists and other restrictions require separate assessment. Neither experiment measured actual payment decisions, employee behavior or business harm; consumer ChatGPT and other AI models were not tested.
Trace screening results to their source records
The study’s separate source-data audit found that SanctionsKit’s stored OFAC file matched the official file’s fingerprint and all 19,416 record identifiers at 17:48 UTC. The 96 study identities also matched in membership and compared identity fields. Retrieval was recorded at 17:11 UTC, before the evaluated API test began.
SanctionsKit monitors active sources against their refresh schedules and flags overdue checks, blocked refreshes and failed imports. The team can use those findings to investigate promptly, while automatic retries handle recoverable acquisition failures. Source dates, identifiers and versioned evidence make those checks traceable to the screening result.
The OFAC integration guide explains how the API and review interface work together. Source status shows scheduled refreshes and the latest successful checks. The evidence documentation describes the retained inputs, source versions and results that support a review.
Read the source-data audit timeline and scope
At the preparation check at 17:10 UTC, SanctionsKit held the 5 October edition: 55 records newly present in the 8 October file were not yet stored, and two removed vessel records remained. The 96 selected study identities agreed with the official reference at both observations.
Product metadata recorded retrieval of the 8 October file at 17:11 UTC, before the API test. At the 17:48 UTC check, the official file hash and all 19,416 record identifiers matched. The research team did not trigger the refresh. OFAC source checks are scheduled every six hours; source timestamps make the interval observable.
This audit compared source versions, record identifiers and selected identity fields. The AI experiment’s supplied-evidence packets were assembled by the researchers. The experiment did not test SanctionsKit’s screening endpoint, fuzzy matching or ownership rules.
Reproduce, cite and contact
Download the complete evidence archive and archive manifest. The archive preserves the original experiment alongside source records needed to inspect it. The original methodology, prompts, company panel, all visible responses and case outcomes are also available separately.
SanctionsKit designed and funded this study and sells sanctions screening software. Nick is the article’s research author. Research enquiries: contact SanctionsKit at support@sanctionskit.com through the contact page, referencing this study. No external peer review or government endorsement is claimed.
Suggested citation: SanctionsKit (2026), “When AI gets sanctions status wrong: a 96-company test,” 8 October 2026, study and evidence archive. Identify the Google observation, original API experiment or API follow-up, and include its condition and denominator when quoting a result.
Version note, 8 October 2026: the original 1,152-response experiment and its v1 files remain unchanged. A separate API follow-up adds 960 scheduled decisions and a v2 evidence archive. The completed Google observation adds 96 searches and a v3 evidence archive. Results are kept separate by experiment, model and condition. The company process recommendation concerns reproducible source checks, retained evidence, human review and employee training.
Sources and references
Authority and technical sources
- OFAC Framework for Compliance Commitments (opens in a new tab)
- OFAC Sanctions List Service (opens in a new tab)
- OFAC Sanctions List Search (current data) (opens in a new tab)
- OFAC list modernization notice: 28 May 2026 (opens in a new tab)
- OFAC list modernization notice: 27 July 2026 (opens in a new tab)
- OFAC list modernization notice: 5 October 2026 (opens in a new tab)