Designing the analytics function: risk-scoring models and their failure modes
Figure 3.1 · Geography
One kickback, five jurisdictions
Each hop is a deliberate secrecy choice — a doctrine, a treaty, a professional silence. The final step is always a legitimate-looking asset.
Source · Schematic based on ICIJ Panama/Pandora Papers narratives
Every revenue authority or FIU eventually reaches the point where the volume of reports, returns and transactional data flowing in vastly exceeds the capacity of human analysts to review case by case, and the institution must decide how to build a risk-scoring or analytics function to triage that volume. I have advised on several such build-outs across African revenue authorities and FIUs, and the recurring lesson is that the technology is rarely the binding constraint — governance, data quality and institutional incentives are.
Start with what the analytics function is actually for. In an FIU, it typically means scoring incoming suspicious transaction reports (STRs) and currency transaction reports against a combination of rule-based red flags (structuring patterns, PEP linkages, high-risk-jurisdiction counterparties, rapid in-and-out movement) and, increasingly, statistical or machine-learning models trained on historical case outcomes to predict which reports are most likely to correspond to genuine underlying criminality. In a revenue authority, it typically means scoring filed returns and third-party data (CRS files, customs declarations, VAT invoicing data) to select cases for audit, again using a mix of rule-based indicators (mismatches between declared income and asset holdings, industry-specific margin anomalies, related-party transaction volumes inconsistent with arm's-length pricing) and predictive models trained on the outcomes of past audits.
The design choice between rule-based and machine-learning approaches is not merely technical; it has real institutional-integrity consequences. Rule-based systems are transparent and auditable, an official facing a legal challenge or an oversight review can explain exactly why a case was flagged, because the rule is legible. Machine-learning models, particularly more opaque architectures, can achieve higher predictive accuracy on held-out historical data but are harder to explain to an affected taxpayer, a court, or a legislative oversight committee, and; critically, are vulnerable to reproducing and amplifying whatever bias exists in the historical case outcomes they were trained on. If past audit selection was itself biased (say, disproportionately targeting small informal-sector traders because they were easier targets, while systematically under-auditing complex multinational structures because they required scarce specialist skills the authority lacked), a model trained on that history will learn and perpetuate exactly that bias, dressed up in the appearance of statistical objectivity. I tell every analytics team I advise: a model is not neutral merely because it is mathematical: it inherits every bias present in its training data, and unexamined deployment of such models in a fragile institutional-trust environment is a serious risk, not a shortcut.
A second, distinct failure mode is overfitting to known typologiesOverfitting to known typologiesA model's weakness in detecting genuinely novel schemes because it was trained only on previously identified and caught patterns. at the expense of detecting genuinely novel schemes. Risk-scoring models are trained on cases that were caught; by construction, they are weaker at flagging the structurally different scheme that hasn't yet been identified as a typology, because there is no positive training example for the model to learn from. This is precisely why FATF guidance and virtually every serious analytics practitioner insist that automated scoring must remain a triage tool that feeds a human analytical review, not a replacement for it, and why a meaningful proportion of an FIU or revenue authority's analytical capacity should remain devoted to open-ended, hypothesis-driven investigation of patterns the models were never trained to look for.
A third failure mode, common in under-resourced institutions, is what I call "scoring without capacityScoring without capacityThe failure mode of generating more correctly flagged high-risk cases than an institution has resources to investigate, creating an accountability exposure." — building a sophisticated risk-scoring system that correctly identifies far more high-risk cases than the institution has capacity to actually investigate, resulting in a growing backlog of correctly flagged but never-actioned cases. This is worse than not scoring at all in one specific respect: it creates a documented institutional record, discoverable in any later audit, mutual evaluation or oversight review, that the authority knew about the risk and failed to act — a serious accountability exposure. The correct design discipline is to size the scoring threshold to actual investigative capacity, accepting a higher score cut-off (and therefore fewer, more confidently flagged cases) rather than generating a flood of flagged cases the institution cannot realistically work.
Data quality is the unglamorous but decisive determinant of whether any of this works. Analytics built on inconsistent taxpayer identification numbers, unreconciled beneficial-ownership data, or STR narrative fields filled in inconsistently by reporting entities will simply produce garbage scores regardless of the sophistication of the modelling technique layered on top. I have seen institutions spend disproportionate resources on machine-learning model development while the underlying registry data, the beneficial-ownership register, the taxpayer master file — remains riddled with duplicate records, stale addresses and unresolved identity-matching problems that no model can compensate for. The sequencing discipline is: fix identity resolutionIdentity resolutionThe data-quality process of reliably matching records referring to the same individual or entity across disparate datasets and registries. and data quality first, deploy transparent rule-based scoringRule-based scoringA risk-scoring approach using explicit, legible indicators (e.g. structuring patterns, PEP status) that can be individually explained and audited. second, and only then layer in more sophisticated predictive modelling once there is a sufficient, well-governed data foundation and enough historical outcome data to train against responsibly.
Finally, governance of the model itself needs the same institutional rigour as governance of any other enforcement power. Every risk-scoring model deployed in a state institution with coercive powers over citizens' financial affairs should have a documented validation methodology, periodic bias and performance audits against demographic and sectoral breakdowns (not just aggregate accuracy), a clear human-override and appeal pathway for affected persons, and an accountable owner within the institution who can explain and defend the model's design and outcomes to oversight bodies, courts, and, where BO or personal data is involved — a data-protection authority. This last point is the subject of the next lesson.
Risk-scoring approach against transparency and predictive power
Institutions with fragile public trust should weight toward transparency even where it costs some predictive nuance.
Key terms
- Rule-based scoring
- A risk-scoring approach using explicit, legible indicators (e.g. structuring patterns, PEP status) that can be individually explained and audited.
- Overfitting to known typologies
- A model's weakness in detecting genuinely novel schemes because it was trained only on previously identified and caught patterns.
- Scoring without capacity
- The failure mode of generating more correctly flagged high-risk cases than an institution has resources to investigate, creating an accountability exposure.
- Identity resolution
- The data-quality process of reliably matching records referring to the same individual or entity across disparate datasets and registries.
Exercise
Design a governance checklist (8-10 items) that a revenue authority should apply before deploying any new risk-scoring model into live case-selection use, covering data quality, bias testing, capacity-sizing and appeal rights.
Mark complete (sign-in) →Sources
Last reviewed 2026-08-01
- 01FATF Guidance on the Use of Digital Identity and analytics in AML/CFT contexts — FATF, 2023.Guidance emphasising human oversight of automated risk-scoring tools.
- 02ATAF technical notes on tax administration risk analytics and audit case selection — African Tax Administration Forum, 2022.
- 03OECD Forum on Tax Administration guidance on advanced analytics for compliance risk management — OECD, 2021.