How Accurate Are Crypto KYT Risk Scoring Tools?

A Score Is Only as Good as the Evidence Behind It

KYTComplianceRisk Scoring
August 15, 20269 min read

Behind every blocked deposit or held withdrawal sits a number: the KYT risk score. That single number decides whether a deposit clears, whether a withdrawal is held, and whether a suspicious transaction report is filed. The question buyers and examiners ask is the same: how accurate is that score, and can the team explain why it landed where it did. This guide breaks down what crypto transaction risk scoring tools measure, the accuracy vs explainability tradeoff, why regulators demand scores they can read, and what to look for when evaluating a tool. For the screening workflow behind these scores, see Phalcon Compliance. This page is part of the KYT Resource Center.

What KYT Risk Scores Measure

A KYT risk score is a summary of the illicit exposure attached to an address or transaction. It is built from two inputs: direct exposure to known illicit counterparties, and indirect exposure that travels through multi-hop fund paths.

Direct exposure is the simpler input. When an address has transacted with a sanctioned entity, a darknet market, a mixer, or an address flagged by a law enforcement seizure, that connection is a direct hit with a label. Indirect exposure is where scoring engines diverge. Funds move in hops, and layering exists to dilute traceable exposure. An engine that ignores multi-hop paths undercounts risk, and one that treats every hop as equally risky overcounts it.

The engines that handle this well quantify multi-hop contamination. Phalcon Compliance reports Exposure Value, the absolute amount of illicit funds that reached the address through traced paths, and Exposure Percentage, that value as a share of total inflows. Together they answer the two questions an examiner asks: how much dirty money touched this address, and how much of its activity is dirty. The first drives a freezing decision. The second drives a de-risking decision.

On top of exposure, the score aggregates Risk Indicators. It surfaces 17 Risk Indicator categories spanning sanctioned entities, darknet markets, mixers, scams, thefts, terrorist financing, child safety, fraud, gambling, and other illicit and high-risk activity. Each indicator is a labeled, attributable signal, and the score is the weighted combination of those indicators plus the exposure figures, not an opaque model output.

Full address risk detail showing score breakdown and exposure figures

Trade Detection Power for Explanation

The accuracy question is not a single axis. There are two kinds of accuracy, and they pull in opposite directions.

The first is detection accuracy: how well the engine catches new, previously unseen laundering patterns. Black-box machine learning models have an edge here. A model trained on graph patterns can flag a structuring or layering sequence it has never seen labeled, because the sequence looks statistically similar to known bad behavior. For emerging typologies, a behavioral model catches what a label-only engine misses.

The second is explainability: whether the team can say, in writing and under examination, why a score is what it is. Label-driven glass-box engines win here. A score built from named Risk Indicators and quantified Exposure Value is explainable by construction. The compliance officer can point to the indicator, the counterparty, and the amount, and write that into a suspicious transaction report. A score driven by a model's internal weights cannot be unpacked the same way, and the gap between the score and the explanation is where examination findings live.

The tradeoff is not resolved by buying a more expensive tool. A glass-box engine is only as current as its labels. If a new mixer launches this month and the label has not propagated, the engine will undercount risk on addresses that already transacted with it. A behavioral model catches the pattern earlier but struggles to explain the catch in a way a regulator accepts. The realistic position is that the strongest scoring layer combines both: label-driven indicators for explainability and behavioral detection for emerging patterns.

Why Regulators Demand Explainable Scores

The FATF risk-based approach obliges a VASP to understand the risk it is taking on, and FinCEN's AML program rule (31 CFR 1022.210) requires the program itself to be documented and available for inspection, which is why per-decision reasoning has to be reconstructable. Regulators do not grade a risk score in isolation. They grade whether the team can defend the decision that score drove. A high score that triggered a freeze has to be justifiable. A low score that cleared a transaction that later turned out to be illicit has to be justifiable too. A score the compliance team cannot explain reads as a program gap, regardless of how accurate the underlying engine is.

This shows up in operational practice. Compliance professionals working with crypto screening tools describe the demand to explain risk scores to regulators as a live operational pressure, not a theoretical preference. The recurring theme across examination cycles is the need to explain scoring decisions to examiners, with examination findings framed as the consequence of scores that cannot be unpacked.

The regulatory frame backs this up. The FATF risk-based approach for virtual asset service providers expects ongoing monitoring and an auditable methodology, not just an outcome. Examiners read an auditable scoring methodology as evidence of a working program. They read a black-box score they cannot audit as a control that exists on paper but cannot be verified, which is worse than a simpler control that can be verified. Accuracy without explainability fails the examination test. Explainability without accuracy fails the detection test.

Risk indicator counterparty graph showing which signals drove the score

Glass-Box Scoring: 17 Risk Indicator categories and Exposure Quantification

A glass-box scoring engine exposes its inputs. Phalcon Compliance builds on this principle. Every risk score it returns is traceable to a named set of 17 Risk Indicator categories and to quantified Exposure Value and Exposure Percentage. The compliance officer reading the score can see which indicators fired, which counterparties triggered them, and how much illicit value flowed through the address.

The 17 Risk Indicator categories cover what examiners expect: sanctioned entities and addresses, darknet markets, mixing services, scams and fraudulent projects, stolen funds, terrorist financing, child exploitation material, gambling, and additional high-risk and illicit behavior. Each indicator is labeled and attributable. The score is an aggregation of named signals, and each signal can be inspected, challenged, and written into a disposition or a regulatory filing.

Exposure Value and Exposure Percentage sit on top of the indicators as the quantification layer. They answer the multi-hop contamination question that binary flags cannot. An address with a small Exposure Percentage but a large Exposure Value is a high-volume address with limited illicit contamination. One with a large Exposure Percentage and a small Exposure Value is a small address dominated by illicit inflows. A glass-box score that reports both lets the team pick the right response and freeze the right funds.

The practical result is that the score is auditable by construction. The team does not need to reconstruct the reasoning after the fact. The reasoning travels with the score.

Behavioral Detection Plus Exposure: Catching New Layering While Staying Explainable

The weakness of a label-only glass-box engine is label latency. New typologies emerge faster than labels propagate, and sophisticated layering is designed to dilute traceable exposure below the threshold where a label-driven score fires. A scoring layer that relies only on historical labels will miss new patterns until the labels catch up, and by then the funds have often moved.

The response is a Behavioral Risk Engine layered on top of the indicator and exposure scoring. The platform includes a Behavioral Risk Engine that detects patterns associated with layering, structuring, and fund movement typologies that may not yet carry a label. It looks at how funds move, not just where they came from, which lets it flag activity that looks like laundering even when no individual counterparty is on a sanctions or illicit list.

This combination matters because it preserves explainability while extending detection reach. The behavioral signal does not replace the indicator and exposure scoring. It adds a layer. When the behavioral engine flags an address, that flag is a Risk Indicator in its own right, with its own label and its own audit trail. The compliance officer still sees a named, attributable signal, not a raw model output.

This is the realistic resolution to the accuracy vs explainability tradeoff. A glass-box engine that pairs labeled indicators and quantified exposure with a behavioral detection layer can catch emerging layering patterns without abandoning the explainability that examination requires.

Evaluating a Scoring Tool for Your Team

When a compliance team evaluates a crypto transaction risk scoring tool, the accuracy question has to be broken into the three things that actually matter under examination. A tool strong on one and weak on the other two will fail at the point where it matters most.

First, can the score be explained. The tool has to return named, attributable Risk Indicators alongside the score, plus quantified exposure figures the team can read into a disposition or a filing. Ask the vendor for a sample score breakdown and check whether the indicators and exposure numbers are visible, labeled, and traceable to counterparties. A score without that breakdown is a score the team will have to defend without evidence.

Second, can the score be audited. The tool has to retain the screening history, the score at the time of screening, and the indicators and exposure figures that drove it. Examiners ask what the score was on the day the transaction was cleared, not what it is today after the address has been relabeled. A tool that does not freeze the historical record cannot answer that question.

Third, can the score be defended. The tool has to support the team's threshold logic, alert routing, and suspicious transaction report workflow. The score is only useful if it plugs into the program, with risk bands that map to clear, hold, and block decisions, and an audit trail that connects the score to the action taken. A high-accuracy score that does not integrate is a dashboard number, not a control.

The accuracy question, asked honestly, is whether the tool gives the team a score it can act on, explain, and defend in that order. It applies the glass-box principle: 17 Risk Indicator categories, Exposure Value and Exposure Percentage, and a Behavioral Risk Engine that extends detection without sacrificing explainability. Explore Phalcon Compliance to see how the scoring layer surfaces the indicators and exposure figures your examination record depends on.

Alert detail showing scoring evidence attached to each flagged transaction

Frequently Asked Questions

Build Real-Time, Automated, and Auditable KYT Compliance Capabilities

Systematically improve virtual asset transaction risk monitoring capabilities, from understanding regulatory obligations to implementing technical architecture.