Adverse Media False Positives: What 99.4% Noise Reduction Looks Like
Ask a compliance or third-party risk team what makes their job hard, and almost nobody says they can’t find enough information. They say the opposite. There is too much of it, most of it doesn’t matter, and separating the two eats the day.
That is the real state of screening in 2026. Detection is largely a solved problem, the tools find things. What they mostly don’t do is decide what deserves your attention. Industry benchmarks put false positive rates across AML and adverse media screening between 85% and 95%1,2, which means the overwhelming majority of what reaches an analyst resolves to nothing at all.
This blog shows what the alternative looks like in practice, using one real-world scale: a portfolio of 150 monitored companies over a year: 95,000 screened records, and 580 events that actually needed a human.
The real cost of a noisy alert queue
Portfolios have grown. Headcount hasn’t. A team that monitored a few hundred entities five years ago now covers thousands, across sanctions, PEPs, watchlists, state ownership, and news in dozens of languages.
When the queue outgrows the team, three things happen, and none of them are good.
- Real risk gets the same attention as noise. An analyst working through hundreds of low-value alerts is not reading any single one carefully. The genuinely serious item is in there, receiving thirty seconds.
- Scope quietly narrows. Teams stop monitoring the long tail of vendors, or fall back to annual reviews, because continuous coverage of everything has become unaffordable in hours.
- Dismissals stop being defensible. Under supervision models like the EU’s AMLA framework, the question is shifting from whether an alert fired to whether it was reasonable for it to fire, and whether the decision to close it was documented3. A queue cleared under time pressure is an audit problem waiting to happen.
None of that is fixed by finding more. It’s fixed by filtering better, before the work reaches a person.
What 95,000 records actually look like
Here is a portfolio of 150 companies across a full year of monitoring, start to finish.
Across those 150 companies, Owlin selects 95,000 candidate records over the year. Irrelevant matches and mistaken entity hits are filtered out, leaving 26.5% of intake as relevant records. Duplicate coverage of the same development is merged, bringing that down to 6%. What remains is scored for materiality, surfacing 580 events as elevated risk: 0.61% of everything screened, a 99.4% reduction in review volume.
One elevated-risk event for every 164 records screened. Put another way: monitoring 150 companies continuously for a year produces around eleven items a week that genuinely need someone. That is the number that decides whether continuous monitoring is something your team can actually sustain.
Importantly, the other 99.4% hasn’t gone anywhere. Nothing is deleted at any stage: rejected records stay retrievable along with the reason they were filtered, merged events keep every source attached, and events that didn’t meet the threshold remain on the entity’s timeline. The funnel decides what comes first, not what exists, so once the elevated 0.61% is worked through, everything behind it is still there to go into as deeply as the situation warrants.
Three steps, three different problems
The reduction isn’t one clever relevance score. It’s three separate steps, each removing a different kind of noise, which is what makes the result explainable rather than a black box.
| Step | What it removes | Result |
|---|---|---|
| Relevance | Wrong entity, incidental mentions, no risk signal | 73.5% of records filtered out |
| Deduplication | The same development reported by many sources | 26.5% of intake becomes 6% |
| Scoring | Real events that don’t move the risk picture | 0.61% flagged as elevated risk |
The distinction that matters most here is the second one. Deduplication is not the same as relevance filtering, and most platforms treat them as one problem. A bankruptcy filing covered by a wire service, three national papers, and a trade publication produces five entirely legitimate, correctly-matched records. None of them is a separate risk event. Merging them into one development with all five sources attached is the difference between reading a story once and reading it five times.
Plus: you only hear when risk changes
On top of those three steps, there’s one more thing worth knowing. Even within the elevated 0.61%, Owlin notifies you based on a change in a company’s risk level rather than on an event simply existing. If a company is already flagged as high risk, being told again that it’s high risk carries no information, so every alert that reaches you says something you didn’t already know.
See what your portfolio looks like
Every portfolio produces a different ratio, depending on entity count, jurisdictions, and risk appetite. Owlin supports major organizations worldwide and has been named a Chartis Category Leader for Adverse Media Monitoring in 2024, 2025, and 2026.
What this actually changes for your team
The 99.4% figure is easy to quote and easy to misread. It doesn’t mean 99.4% of the data was thrown away, nothing is deleted, and everything stays searchable. It means 99.4% of it never had to compete for an analyst’s attention.
In practice, that changes three things.
- The queue becomes finishable. 580 items across a year is a workload a team can genuinely clear, with enough time on each to make a considered call. That is a different job from triaging tens of thousands, and once it’s clear, the rest of the data is still there to go deeper into.
- Coverage stops being a trade-off. You can monitor the entire portfolio continuously rather than choosing which third parties are worth watching. Full coverage and a manageable queue stop being mutually exclusive.
- Each item arrives ready to work. Because every stage keeps its evidence, an elevated-risk event comes with its sources, its timeline, and the reasoning behind its score. Analysts spend their time deciding, not assembling.
The comparison worth making isn’t against another vendor’s feature list. It’s against the industry’s own baseline: at an 85–95% false positive rate, most of a compliance team’s capacity is spent on findings that go nowhere1.
Filtering this aggressively only works if it’s explainable
A platform that keeps 99.4% of what it ingests out of your queue is making a very large number of decisions on your behalf. If it can’t show why, you haven’t solved alert fatigue, you’ve replaced it with something harder to defend.
This is why Owlin’s risk scores come with reasoning attached rather than as a number on its own. Every merged source, and every score can be reconstructed. When a regulator, an auditor, or your own risk committee asks why an event was escalated or closed, the answer exists and is traceable to source.
Evidence first, then AI. Explainable by design. Humans making the decisions. Aggressive filtering is only an asset if all three hold.
Where this applies
The ratio isn’t specific to one use case. The same pipeline runs across:
- Third-party risk and vendor risk, where portfolio size makes manual review structurally impossible
- Merchant risk for payment providers screening at onboarding and continuously afterwards
- Supply chain and ESG obligations under CSDDD, the German Supply Chain Act, and ESG frameworks
- Adverse media monitoring for risk and compliance teams
The pressure is the same everywhere: too little coverage is a compliance failure, too much is an operational one. The ratio is what resolves it.
See it on a company you care about
Run a free check and see the risk events, evidence, and scoring for a company you already know.
Frequently asked questions
What is a typical false positive rate in adverse media screening?
Industry benchmarks place false positive rates across AML and adverse media screening between 85% and 95%1,2. Adverse media tends to sit at the higher end because the source material is unstructured and effectively unlimited, unlike finite sanctions or PEP lists.
How does Owlin reduce false positives?
Through four sequential filters. A relevance stage rejects records that fail entity verification or carry no risk signal, 73.5% of everything screened. A deduplication stage merges related coverage into single events, cutting what remains to 6% of intake. Scoring identifies which events are materially elevated. A final filter notifies only when a company’s risk level actually changes.
What percentage of screened records reach an analyst?
0.61%. Across a portfolio of 150 monitored companies over a year, 95,000 screened records produced 580 elevated-risk events requiring analyst attention, a 99.4% reduction in review volume, or one event per 164 records screened.
Does filtering that aggressively risk missing real risk?
No records are deleted. Rejected records stay traceable, and events that aren’t elevated remain on the entity’s timeline and stay searchable. The reduction is in what competes for attention, not in what is retained or findable.
Sources
- Facctum. (2026). AML false positive rates 2026 report: Statistics, costs and remediation.
- Sigma360. (2026). Adverse media false positives: A complete 2026 guide.
- Biometric Update. (2026). Under AMLA, 95% false positives become a regulator’s problem.