AML/CTF obligations for Tranche 2 businesses are now in force. Since 1 July 2026 Get compliant
AML False Positives Are a Data Quality Problem. The Fix Starts at Intake.

AML False Positives Are a Data Quality Problem. The Fix Starts at Intake.

Industry estimates put AML alert false positive rates at 90 to 95 percent. The cause sits more in the data going into the screening engine than in the engine itself.

AML/CTF Compliance 10 May 2026 10 min read AML Guard

For most reporting entities, especially those preparing to come under regulation for the first time on 1 July 2026, the question of where AML customer data originates is treated as a back-office detail. That is a mistake. The point at which a customer's name, date of birth, address, and source of funds enters the compliance record, and the order in which the platform verifies and screens that data, together shape how many false positives, how many enhanced due diligence escalations, and how much officer time will be spent investigating things that turn out to be nothing.

A quick orientation for readers new to the language. Customer due diligence (CDD) is the regulator-required identity and risk-assessment work before a reporting entity transacts with a customer. Enhanced customer due diligence (ECDD) is the additional checks that kick in for higher-risk customers: for example, politically exposed persons, unusually large transactions, or customers from high-risk jurisdictions. Identity verification (IDV) is the document-and-biometric step inside CDD where the customer photographs their ID and takes a selfie. A politically exposed person (PEP) is someone in a senior public position whose accounts carry higher risk of corruption-linked money flows. With those terms in hand, the rest of this article is about how the data feeding all of that gets captured, and what happens to false positive rates when the architecture is wrong.

Why Do AML Systems Generate So Many False Positives?

The AML false positive rate sits between 85 and 95 percent across most published industry research, and the largest single contributor sits upstream of the screening tool itself.

PwC's analysis puts the rate at 90 to 95 percent. LexisNexis Risk Solutions describes institutional false positive rates of 95 percent or more as not unusual. Industry commentators routinely refer to a 95 percent benchmark as the industry's dirty secret. Global AML compliance costs are estimated at over USD 274 billion annually, with much of that consumed handling alerts that turn out to be noise.

Several mechanical causes contribute. Name-matching algorithms that trip on common names. Watchlists with insufficient secondary identifiers. Static thresholds that have not been recalibrated. Missing entity resolution that fails to merge duplicate records. Each is real and well-documented.

But across recent industry analysis: Ondato, LexisNexis Risk Solutions, Kroll, FATF guidance, the Wolfsberg Group's published sanctions screening guidance, the same factor sits at the top of the root cause list: the quality of the input data the system has to work with.

The principle: the engine searches a customer record against watchlists, sanctions data, and PEP information. The engine can only work with the identifiers it is given. If the customer record itself is wrong: a misspelled surname, a transposed digit in a date of birth, an incomplete address, the engine returns the matches the data tells it to return.

What Regulators and Standard-Setters Actually Say About Data Quality

The Wolfsberg Group, FATF, and AUSTRAC each treat data integrity at intake as a foundational compliance control, not a hygiene issue.

The Wolfsberg Group's Guidance on Sanctions Screening (January 2019) lists data integrity as one of the named pillars of an effective screening programme. The guidance is explicit: a screening programme must include validation of the integrity, accuracy and quality of data so that accurate and complete data flows through the monitoring and filtering systems. Data integrity is treated alongside governance, technology, alert handling, and independent testing, not as a sub-bullet under any of them.

The Financial Action Task Force, in its Guidance on Digital Identity (2020), goes further. Recommendation 10 of the FATF Standards requires regulated entities to identify and verify customer identities using reliable, independent source documents, data, or information. The 2020 guidance explicitly recognises that reliable digital identity systems can minimise weaknesses in human control measures. That phrase is doing real work. FATF is not saying digital systems are inherently better than human-administered ones. It is saying that human handling of customer data is a known weakness, and that the digital channel is expected to reduce it.

AUSTRAC's customer identification and verification guidance illustrates the same point with a published worked example, often referred to as the SmallBank/ShopCo example. A mutual bank is carrying out applicable customer identification procedures on a small retail business with three directors. A staff member receives the company information, verifies it against the documents provided, and enters it into the bank's system, but accidentally misses one of the three directors. Transaction-monitoring alerts later fire and the bank is forced into enhanced customer due diligence; only then does the link to the omitted director come into view.

This is a regulator-published worked example rather than a named enforcement action, but that is precisely why it is useful. AUSTRAC chose it to illustrate a representative failure mode, not a one-off. The structural lesson sits at the architecture level: no IDV step can recover a person who was never entered into the data path to begin with.

A broader real-world version of the same problem appeared in the United Kingdom in 2025. On 8 July 2025, the Financial Conduct Authority announced a £21.1 million fine against Monzo Bank for inadequate anti-financial-crime systems and controls. Among the specific FCA findings: Monzo had onboarded customers on the basis of limited, and in some cases obviously implausible, information: including customers using well-known London landmarks (Buckingham Palace, 10 Downing Street, even Monzo's own headquarters) as their residential addresses. The FCA noted that Monzo's onboarding validated only that an input resembled a UK postcode format, not whether the address was real. Different fact pattern from ShopCo, same architectural lesson: the integrity of customer data at intake is a financial-crime-controls issue, not housekeeping. A platform that accepts Buckingham Palace as a residential address generates downstream alerts whose true cause is upstream data, not a flawed screening engine.

How Transcribed Customer Data Becomes Tomorrow's False Positives

When customer data is typed into a system by anyone other than the customer, every keystroke is a chance for a transcription error to enter the compliance record.

The category labels for these errors are familiar. Misheard surnames. Misspelled middle names. Transposed digits in dates of birth. Addresses captured for marketing purposes: no unit number, abbreviated street type, rather than for identity verification. Beneficial owners missed because the company structure was simplified into a contact record. Source of funds left blank because the agent did not think to ask in the discovery call.

Each of these, individually, looks small. Run them through a sanctions or PEP screening engine and they compound into the 95 percent figure.

A misspelled surname produces fuzzy-match alerts against unrelated sanctioned individuals. A wrong date of birth fails to distinguish the customer from a same-name PEP. An incomplete address removes the secondary identifier Wolfsberg specifically describes as the key element separating a true match from a false one. A missed beneficial owner means the customer profile is structurally incomplete, so when ongoing monitoring fires, the alert lands without context and the officer has nowhere to look first.

Kroll, in its published guidance for compliance officers, puts it in compact terms: if the data captured at onboarding is inaccurate, out of date, or missing, that wave of bad data flows through subsequent ongoing monitoring processes, making it less reliable for every later use. Bad data at intake is not contained to intake. It compounds through every screening cycle, every periodic review, every transaction monitoring rule, for the life of the customer relationship.

The architecture point is straightforward. The further away the data capture is from the customer, the more transcription steps sit in the path, and the more chances small errors have to enter the record. A model where a customer's details are first captured by a salesperson, retyped into a sales contact record, retyped again or pushed into a compliance module, and then handed to the screening engine has up to three places where small errors compound into noisy alerts.

The architecture rule: every transcription step between the customer and the compliance record is a chance for the data to degrade. Fewer transcription steps means cleaner input and a lower false positive rate for the screening engine to manage.

Doesn't Identity Verification Catch These Errors?

Identity verification corrects a narrow band of identifying fields. It does not verify or correct most of the substantive AML/CTF data that drives risk scoring, ECDD escalation, or beneficial ownership decisions.

This is the obvious rebuttal to the data-quality argument: surely IDV catches the typos? The honest answer is that IDV catches a small subset of fields on the natural person being verified, and does nothing for the rest.

What IDV verifies, in a typical Australian DVS-backed flow:

FieldVerified by IDV?
First name and last name (against the document)Yes
Date of birthYes
Document number and expiryYes
Liveness and biometric face matchYes
AddressVerified only where the document carries it, driver licences print residential address, passports do not

What IDV does not verify, capture, or correct, and these are typically the larger contributors to false positive rates:

FieldIDV touches it?Why it matters for false positives
Source of funds narrativeNoRisk score input; ECDD trigger
Source of wealthNoRisk score input
Email and mobileNoIdentity correlation for ongoing monitoring
Citizenship and country of birthNoJurisdictional risk input
Occupation and employerNoPEP-adjacent risk signal
Relationship type, buyer, seller, tenant, executorNoDetermines the entire CDD path
Self-declared PEP statusNoRisk score input
Beneficial ownership structureNoThe AUSTRAC ShopCo failure mode
Each beneficial owner's identityOnly if each owner independently completes IDVWhy entity CDD is harder than individual CDD

The AUSTRAC SmallBank/ShopCo worked example sharpens here. The third director was missed at the data-entry step. No IDV step can recover a person who was never entered to be verified. IDV verifies a person who is being checked; it cannot conjure a person who was structurally omitted from the data path.

The same logic applies to source of funds. If a sales agent recorded source of funds as "savings" in a CRM contact note based on a discovery conversation, no IDV step will return to the customer and reconcile that against what the customer would have written if asked directly with the categorisation in front of them. The compliance record carries the agent's paraphrase forward.

The rebuttal that "IDV catches the errors" is, on inspection, a rebuttal about a small number of fields on a small number of people. It is not a rebuttal about the substantive AML/CTF data that drives the false positive rate.

Why Client-Entered Intake Is Structurally Different

When the client enters their own details directly into a compliance intake portal, the transcription error path largely disappears across every field, not just the ones IDV later verifies.

Several things change when the customer enters the data themselves, and they compound.

The client reads directly off their identity document. There is no intermediate listening or typing step. The name spelled on the passport or driver licence is the name typed into the form, because the client is looking at the document while doing it.

The client has the strongest incentive to enter their own details correctly. If they make a mistake, their own identity verification will fail and they will be the one inconvenienced. That self-interest is a quality control mechanism no third-party transcription path replicates.

The client confirms before submission. Most well-designed intake portals show the captured data back for review and require acknowledgement that it is correct. That acknowledgement is a real legal artifact: the client is attesting to their own data, which has evidentiary weight if a question is later raised about whether the information was correctly captured.

Source of funds is captured directly from the customer, against the structured categorisation a compliance program needs. The narrative is the customer's, not a sales agent's reconstruction of a discovery conversation.

For entity matters, beneficial owner invitations are generated automatically from the initial submission, and each beneficial owner enters their own details directly. The chain of transcription that would otherwise produce a missed director (the SmallBank/ShopCo failure mode) is removed by design.

None of this means the screening engine matters less. It matters as much as it always did. But the engine is now operating on a much cleaner input set, and the false positive rate it produces drops accordingly.

When Should Screening Fire? Before or After Identity Verification?

For full onboarding screening, match quality is generally stronger once core identity data has been verified, because the engine has more reliable identifiers to disambiguate true and false matches. Pre-verification screening should be tightly bounded and not treated as full CDD clearance.

This is the second false positive amplifier hidden in market positioning. Several platforms market a pre-emptive PEP and sanctions screen, variously described as "instant PEP check," "real-time name screening," or "screen at intake", as a feature. The pitch is that catching risk early is better than catching it later. Operationally, the trade-off is more complicated.

Regulators require that PEP and sanctions status be established before designated services are provided. They do not generally mandate that the screening run occur after IDV specifically. The honest claim is operational, not regulatory: the data the screening engine needs to disambiguate a true match from a false one (date of birth, document number, full middle name where present, address) is precisely the data identity verification produces. Running screening before IDV means running it without those secondary identifiers, and the result is structurally noisier output.

What that looks like in practice, field by field:

FieldAvailable pre-IDV?Verified by IDV?Effect of running screening pre-IDV
First and last nameYes (typed)YesEngine searches on a typed name with no confirmation it matches a real document
Date of birthYes (typed)YesA transposed digit either misses a true match or causes a false collision against a same-name PEP
Document numberNoYesScreening has zero document context
AddressYes (typed)SometimesWolfsberg's named secondary disambiguator is unverified or missing

The output of a pre-IDV screen on the same customer is structurally noisier than the output of a post-IDV screen. It is the textbook example of running a control on data whose integrity has not been validated, exactly the failure Wolfsberg's data integrity pillar exists to prevent.

There is a legitimate exception. Some workflows, auction pre-registration is the clearest example: need a sanctions-only check on a registrant before allowing them to bid, well before a full CDD case is opened on the eventual winner. A well-designed pre-registration check is bounded, sanctions-only (not PEP, not adverse media), and explicitly demarcated as not CDD approval. That is a proper bounded exception with guardrails. It is not the same as marketing pre-emptive PEP screening as an everyday onboarding feature. The distinction matters: bounded exceptions with explicit guardrails do not weaken the data integrity pillar. Marketing pre-IDV screening as a routine part of onboarding does.

The screening rule: where match quality matters, screen on verified data. Where a workflow legitimately needs a pre-IDV check (auction pre-registration, walk-in event entry), bound it tightly, scope it to sanctions only, and make explicit that it is not CDD.

What This Changes for Tranche 2 Reporting Entities

For real estate agencies, conveyancers, and property developers coming under regulation on 1 July 2026, day-one decisions about where customer data originates and when it gets screened will determine three years of operational pain or efficiency.

Most Tranche 2 entities will be implementing AML/CTF compliance for the first time. Many will assume: reasonably, given how the market is positioned: that the choice of compliance platform is mostly about features, pricing, and integration into existing systems. The data quality question rarely makes the shortlist. The screening order question rarely appears at all.

Both should. Three years from now, the agencies that calibrate their alert volumes to AUSTRAC's expectations will find themselves either in a sustainable operating model or in a permanent backlog. The variables driving that outcome are not primarily which screening engine they chose. They are whether their customer data is being captured by the customer or transcribed by their own staff into a sales system that was never designed for compliance integrity, and whether their screening engine is given the cleanest input set the platform can deliver.

This is also a defensibility point. Under AUSTRAC's reformed AML/CTF framework, the new initial customer due diligence obligations require entities to have collected and verified customer information using reliable and independent sources before designated services are provided. An intake architecture that originates customer data from the customer, with self-attestation and document-backed verification, then runs screening on the verified data, sits naturally inside that requirement. An architecture that relies on agency staff to transcribe customer details from a discovery conversation, runs screening on that typed data, and treats IDV as a parallel step that happens to catch some of the typos, sits less naturally.

What a Low-False-Positive Intake Architecture Looks Like

The decisions that drive false positive rates down are not primarily about software features. They are about who enters the data, when verification happens, when screening fires, and how cleanly verified data flows into the compliance record.

A stronger intake architecture has the following characteristics:

Customer-originated intake The customer receives a secure intake link and enters their own details directly: full legal name as it appears on identity documents, date of birth, residential address, and any declarations relevant to the matter. Self-attestation as a captured artifact Before submission, the customer confirms the data is correct. The acknowledgement is recorded as part of the case file, creating a defensible record of the customer's own attestation, separate from any later transcription. Source of funds captured directly from the customer The customer answers source of funds questions in the intake flow, in their own words, against the categorisation shown to them. The compliance record receives the customer's narrative, not a sales-conversation paraphrase. Document-backed verification linked in the same flow Identity verification using government-issued documents and biometric matching runs as part of the same client session. The data the customer enters is verified against the data on the document. Verified data captured against the original submission, not over it When IDV completes, the verified name and date of birth are recorded against the customer's original intake submission. The customer's self-attested record is preserved as an immutable audit artifact; the verified record sits alongside it. Any divergence between the two is recorded and inspectable. Screening uses the verified data Sanctions, PEP, and adverse media screening uses the IDV-verified name and date of birth, giving the engine every secondary identifier the document carries, exactly the disambiguation context Wolfsberg's data integrity principle calls for. Beneficial owner self-entry for entity matters Each beneficial owner receives their own portal link and enters their own details directly. The compliance record captures each beneficial owner from source, with the same verified-data-write-back applied to each. Bounded pre-IDV checks where legitimately needed Where a workflow genuinely requires a pre-verification check: auction pre-registration is the clearest example: the check is sanctions-only, time-limited, and explicitly demarcated as not CDD approval. One transcription-free path from customer to compliance record From the customer's keystroke to the screening engine input, no manual re-keying step is required. The compliance officer reviews and decides on a substantially complete, audit-ready case rather than building one from fragments.

This is not the only architecture that can produce low false positive rates. But it is one in which the structural transcription risk is removed by design rather than mitigated by training, and the screening engine works with the cleanest data the platform can deliver. Mitigation through training and process is fragile. Removal by design is durable.

The Quality of Your Alerts Is the Quality of Your Data

The 95 percent false positive figure is at least as much a statement about upstream data and workflow design as it is about the screening engine. Engine quality matters. But cleaner intake data and better sequencing materially improve the signal the engine has to work with, and that is the lever most under-utilised in current Tranche 2 platform decisions.

Wolfsberg, FATF, and AUSTRAC all converge on the same point: data integrity at intake is foundational, not incidental. Monzo's £21 million fine illustrates what happens when that principle is treated as housekeeping rather than as a financial-crime control. The AUSTRAC SmallBank/ShopCo worked example illustrates the same lesson at a more granular level. Both point in the same direction: any system that runs on transcribed customer data will produce more noise than a system running on customer-entered data, regardless of how the platform is branded.

For Tranche 2 entities preparing for 1 July 2026, the questions worth asking of any AML platform are not "how good is the screening?" That part is broadly comparable across credible options. The better questions are:

The honest answers will tell you, with surprising accuracy, what your false positive rate, your enhanced due diligence escalation rate, and your compliance officer workload will look like in twelve months' time.

Read more: How the AML Guard Client Intake Portal removes the transcription step: agency-branded secure link, customer self-entry, declarations and source of funds in one flow, identity verification linked, and beneficial owner invitations generated automatically. → /insights/client-intake-portal-aml-onboarding.html

Ready to see AML Guard in action?

Book a demo and see how AML Guard’s client intake portal removes the transcription path that drives false positives, and how the platform sequences screening on verified data.

Book a Demo

Related Reading

The Client Intake Portal: How Smart Agencies Will Onboard Buyers and Sellers Under Tranche 2
Built-In CRM Compliance Sounds Convenient. Here Are Five Questions to Ask.
Three CRM AML Compliance Architectures Emerging Ahead of Tranche 2
When Does AML Compliance Actually Start in a Property Transaction?

This article is for general information purposes only and does not constitute legal advice. Firms should obtain independent professional advice on their specific AML/CTF obligations.

Last reviewed: May 2026.