Email Scam Checker Email Scam Checker
13 min read

How Phishing Detection Accuracy Is Actually Measured

A vendor's accuracy number is measured on a dataset you have never seen. Here is how to read it, and what we found when we measured our own detector.

How Phishing Detection Accuracy Is Actually Measured
P

Pavel Demidovich

Developer and Founder of Email Scam Checker

A phishing detector's accuracy is a property of the dataset it was measured on, not a property of the detector. One of the most important numbers for real-world usability is its false-positive rate on legitimate mail, reported with a confidence interval. Phishing is rare in a real inbox; wrongly flagged ordinary mail is not.

  • Accuracy is one number covering four different outcomes, so it hides which mistake a detector makes.
  • At realistic low phishing prevalence, a detector's alert list is mostly its own false positives - even at a 3.4% rate.
  • Many benchmark datasets are artificially balanced and historical. A real inbox is neither, which is why scores do not carry over.
  • A rate without a confidence interval is a point estimate pretending to be a measurement.
  • We built the measurement harness for our own model, and it caught a failure our test suite had been agreeing with for months.

What an accuracy number actually measures

A phishing detector's accuracy is the share of everything it was shown that it sorted correctly. Put that way it sounds like a complete description of quality. It is not, because a single number is standing in for four different outcomes: phishing correctly flagged, phishing missed, legitimate mail correctly left alone, and legitimate mail wrongly flagged.

The four outcomes form what is called a confusion matrix, and every rate you have ever seen quoted for a detector is some ratio drawn from it. Two cells hold what the detector got right: phishing it caught, and legitimate mail it left alone. Two hold what it got wrong: phishing it missed, and legitimate mail it wrongly flagged, which is the false positive. Reading a confusion matrix is mostly knowing which cell you are standing in.

Accuracy draws from all four cells, precision and recall from two each. That is the whole reason the same detector can be described as 98% accurate and 3% precise without either statement being a lie.

If you only remember one thing from this article, make it this: accuracy is the weakest of the four rates, because it is the only one that improves when the problem gets easier to ignore. A detector that never flags anything scores very well on accuracy if phishing is rare, and it protects nobody.

Accuracy vs precision is not a matter of one being a stricter version of the other. Accuracy summarises a detector's whole run; precision summarises only the alerts it raised. That is why a tool can look strong on one and weak on the other, and why the pair is worth separating before trusting any headline figure.

Metric The question it answers What the value tells you What it hides
Accuracy How much of everything was sorted correctly? Almost nothing on its own Which class it fails on. A detector that flags nothing scores the legitimate-mail rate
Precision Of the emails flagged, how many were really phishing? Its alerts are worth reading Whether it also missed a great deal of phishing
Recall Of the real phishing, how much did it catch? It rarely misses an attack How much ordinary mail it flagged to get there
False-positive rate Of the legitimate mail, how much was wrongly flagged? The tool is unusable and will be switched off Nothing. This is the honest number

Notice the asymmetry in that last column. Precision and recall each hide the other. Accuracy hides both. The false-positive rate is different in kind: it is measured against one class only, mail that is known to be fine, and it counts how often the detector said otherwise. That isolates a failure mode worth isolating, though on its own it says nothing about the phishing the detector missed.

Calculated, the false-positive rate is false positives divided by all legitimate mail - the legitimate column of the confusion matrix. On our own 350-email legitimate split that is 12 wrong flags out of 350, or 3.4%, a number that never refers to the phishing side at all.

The scale of the problem is what makes this more than a technicality. The Anti-Phishing Working Group recorded 971,181 phishing attacks in the first quarter of 2026, up 13.8% from 853,244 the quarter before. The FBI's Internet Crime Complaint Center logged 191,561 phishing and spoofing complaints in 2025, the most-reported crime type that year, with reported losses of $215.8 million.

Phishing statistics at that scale describe the threat environment. They say nothing about whether any given detector is good at its job, and that is a separate question needing different numbers.

For the broader picture of what the detector is up against, our red-flag guide covers the individual warning signs; this article is about how anyone knows whether a machine is finding them.

Why a detector can be accurate and still useless

The base rate is what breaks accuracy, and it breaks it in a way that is invisible unless you run the arithmetic. Phishing is rare relative to ordinary mail. Any detector's mistakes therefore land overwhelmingly on the legitimate side, because that is where the volume is, and a rate like 3.4% produces a large absolute number when the pool is large.

Here is the worked version, using the false-positive rate and recall we actually measured on our own model - 3.4% and 99.4%. The inbox below is an illustrative assumption, not a measurement: 10,000 emails at a phishing prevalence of 0.1%.

In an inbox of 10,000 emails, 10 phishing Count
Real phishing caught (99.4% of 10) 10
Legitimate emails wrongly flagged (3.4% of 9,990) 340
Total alerts raised 350
Alerts that were real phishing (precision) 2.9%
Overall accuracy of the same run 96.6%

Read those last two rows together. The detector is 96.6% accurate and roughly 97 of every 100 things it tells you are wrong. Both numbers are correct. Neither is misleading on its own; the accuracy figure is misleading only when it is presented as the summary of how well the tool works.

The base rate fallacy is the error of judging a detector by how often it is right, while ignoring how rare the thing it is looking for actually is. It is not a flaw in the detector, only arithmetic following from a rare positive class.

No detector controls the prevalence it meets, but it does control the trade-off between catching phishing and flagging ordinary mail - which is why the false-positive rate belongs next to recall, rather than accuracy standing in for both.

There is a second, subtler consequence. A user who receives 340 false alarms a week does not calmly triage them. They stop reading the warnings, or they switch the tool off. In that sense a mediocre false-positive rate does not merely annoy; it destroys the recall the detector was proud of, by training the person to ignore it.

How we measured our own detector

We ship a detector, so this is not a critique written from the sidelines. When we built AI Deep Scan, we also built the harness that measures it. Publishing the result was a deliberate choice rather than a marketing one: the number is less flattering than the one already on the model card, and that is the point.

Our detection model is a DistilBERT classifier that runs entirely on your device. Evaluating it means asking a narrow question: on emails whose correct answer is already known, what does it actually decide? Nothing about that is possible to eyeball, which is why the harness exists as a separate tool outside the shipped extension.

We measured 350 phishing and 350 legitimate emails from a public benchmark, drawn with a fixed seed so the run is reproducible, scoring the model exactly as the extension loads it. Every rate is reported with a 95% Wilson confidence interval, because a point estimate over a few hundred samples reads as far more precise than it is.

The results below are for the decision rule we actually ship, on capped input.

Measured on 350 + 350 emails Result 95% confidence interval
Recall (phishing caught) 99.4% 97.9% - 99.8%
Precision (alerts that were real) 96.7% 94.3% - 98.1%
False-positive rate on legitimate mail 3.4% 2.0% - 5.9%
Missed phishing (false-negative rate) 0.6% 0.2% - 2.1%

Look at the interval on the false-positive rate: 2.0% to 5.9%. That spread is the honest form of the number. At this sample size the harness cannot tell a 2% detector from a 6% one, and a page that printed "3.4%" alone would be overstating what we measured. The interval is not a hedge; it is the actual finding.

What the harness found, and what it caught us doing

The most valuable thing the harness produced was not a rate. It was a bug report.

A previous version of this feature shipped a scorer that returned zero for every email it had ever scanned. It passed its tests, because the tests had been written to agree with the scorer. Nothing in the shipping code was capable of noticing, and no number the product displayed could have revealed it, because the product only displays what the scorer says.

A number produced by the machine is only worth having if something outside the machine checks it. That is the argument for measuring a detector independently of itself, and the check has to be built before the number is believed, not after.

The harness also surfaced something about what our model's output actually is. On 95% of samples the winning class scored above 0.99 with every other class near zero. The model emits a decision, not a calibrated probability, which means the confidence percentage a user sees is not a confidence in any statistical sense. That finding changed how we describe the feature, and we would not have found it by reading the model card.

Why benchmark scores do not transfer to a real inbox

What a benchmark reports is model accuracy measured on its own corpus. It does not predict what the detector will do to your mail, and the gap is not small. Three properties of benchmark datasets push reported accuracy above what any inbox will show.

Benchmarks are balanced; inboxes are not. Test sets typically hold roughly equal numbers of phishing and legitimate messages. Real mail does not, which is exactly the imbalance the worked example above exploits. When a detector meets the real ratio, its precision falls, because the legitimate class it errs on is now vastly larger. Its accuracy barely moves.

Benchmarks are old and narrow. A critical assessment of the field published as E-PhishGen found that most prior work relied on datasets that are not representative of current phishing and mostly cover English. It opens by noting the contradiction directly: despite abundant work reporting near-perfect accuracy, countering malicious email remains unsolved. The datasets and the deployment are describing different worlds.

Same-dataset testing flatters the model. Measuring a detector on data it was trained on lets it learn the shape of that corpus rather than the shape of phishing - its vocabulary, its date patterns, its header quirks. Measured on a corpus it has never seen, performance can drop substantially. This is cross-corpus transfer, and it is why a benchmark figure from a random split should be read as an optimistic estimate under favourable conditions rather than a forecast.

An independent comparison of open-source email filters, published in International Journal of Information Security, makes the same point with a real measurement. Three mature filters were run against the SpamAssassin Public Corpus - 3,000 messages, of which 500 were malicious - with no vendor setting its own test conditions.

Filter Missed of 500 malicious Legitimate flagged of 2,500 Alert precision
SpamAssassin 33 113 80.5%
Rspamd 14 27 94.7%
DSPAM 5 25 95.2%

Convert those miss counts into accuracy and the three look comparable: 95.1%, 98.6% and 99.0%. A buyer skimming that would call it a near-tie and pick on price. Then read the false-positive column. SpamAssassin wrongly flagged four and a half times as much legitimate mail as either competitor, and one in five of its alerts was wrong.

Same corpus, same three tools, and the two columns support opposite conclusions - which is what happens whenever a summary metric and the false-positive rate are allowed to disagree in public.

What is being described Benchmark evaluation Your inbox
Class balance Roughly equal phishing and legitimate Phishing is a small fraction of the total
The phishing in it A fixed historical collection, usually English Whatever was sent this week, in your language
What accuracy tells you How well it learned this corpus Very little, because of the balance above
What the result predicts Capability, under favourable conditions The false-positive rate on representative legitimate mail transfers more directly

None of this makes benchmark evaluation worthless. It is the only way to compare two detectors at all. It simply means the figure answers a question about a model, and the question a user has is about their mailbox - and those two questions do not have the same answer.

Is a bigger dataset the answer?

Not by itself, and the reason is worth stating plainly because it is the intuitive fix everyone reaches for. A larger corpus of the same kind of data makes a measured rate more precise without making it more relevant. You get a tighter confidence interval around a number that still does not describe an inbox.

What would fix the relevance problem is a benchmark built from real, current, multilingual mail at realistic prevalence - and that runs straight into a wall that has nothing to do with engineering. Real, current inbox data would be the most representative corpus, and it cannot be published, because publishing it would mean publishing other people's email.

The E-PhishGen authors address the same constraint by generating synthetic emails instead of collecting real ones. That buys realism in format and language, and gives up authenticity in provenance.

So the field is left with a trade-off it has not resolved, and any accuracy figure you read inherits that unresolved state. That is not a reason to distrust every number. It is a reason to ask which of the two things a number is measuring before you act on it.

The practical version for a user is simpler than the research version. A detector's benchmark score tells you roughly whether it is in the right league. Its false-positive rate, on legitimate mail, with an interval attached, tells you whether you will keep it installed past the first month.

How to audit an accuracy claim in four questions

Any accuracy figure can be audited with four questions, and the answers take a paragraph to give if they exist at all. This is the part that applies whether you are comparing tools or reading a vendor page, and it is short enough to use.

  • Which dataset? A named, public corpus can be checked by someone else. "Our internal testing" cannot, and a vendor who will not name the data is asking you to take the result on trust.
  • What was the class balance? If the test set was balanced and your inbox is not, the accuracy figure does not describe your situation. It is one of the fastest ways to tell whether a reported figure is relevant to your mail.
  • What is the false-positive rate on legitimate mail? Not accuracy, not detection rate. If a page leads with 99.9% and never states this, it is not telling you the number that decides whether the tool is usable.
  • Does that rate come with a confidence interval? Without one, you cannot tell a real difference between two tools from sampling noise - and if the sample is small, neither can the vendor.

A fifth question is worth adding when the tool runs rules as well as a model, or combines several signals: which decision rule was measured, and under what threshold. We report our results for the rule we ship, because a rate for a scorer the user never sees is a number about nothing.

The point of the audit is not to catch anyone lying. Most published figures are honest and were measured in good faith on the data available. The problem is structural: an honest number measured on a balanced historical corpus is still a poor guide to an unbalanced present one, and the gap between the two is large enough to change your decision.

For our own part, the numbers we can stand behind are narrow, and we try to keep them narrow. Our 28 heuristic checks either fire or they do not, and we publish no accuracy rate for them, because we have never measured them the way we measured the model.

Where we have measured - 350 phishing and 350 legitimate emails, capped input, the shipped decision rule - we publish the interval alongside the point estimate, including when it is unflattering.

Final verdict - phishing detection accuracy

Treat every accuracy figure as a property of its dataset, and ask for the false-positive rate on legitimate mail before believing it. If a claim cannot name its dataset, state its class balance, and give that rate with an interval, it is hard to judge how much weight to put on it.

Once an accuracy figure has been read this way, the decision left is architectural rather than statistical: whether detection runs on your own machine or on someone else's server. On-device and cloud-based detection can both be measured honestly - where they differ is what happens to your mail in the process. The interval above is what it looks like when a tool tells you the truth about itself.

Frequently asked questions

Accuracy is the share of emails a detector sorted correctly out of everything it was shown, counting phishing caught and legitimate mail left alone. It is a single number covering two very different kinds of mistake, so it cannot tell you which one a detector actually makes. Precision and recall separate them.

Recall is the share of real phishing a detector catches. Precision is the share of its alerts that turn out to be real phishing. A detector can score high on either alone: catch every phish and flood the user with false alarms, or raise only a handful of alerts and miss most attacks. Precision is the one that reflects how much of the user's attention the tool wastes.

A false positive is a legitimate email a detector flags as phishing. The rate matters more than accuracy because phishing is rare in a real inbox, so a detector's alert list is dominated by its mistakes on ordinary mail. A 3.4% false-positive rate applied to a thousand legitimate emails a week is 34 wrongly flagged messages.

Enough that the error bar is smaller than the claim being made. On 350 samples per class, our own measured false-positive rate of 3.4% carried a 95% confidence interval of 2.0% to 5.9%. A point estimate over a few hundred samples reads as far more precise than it is, which is why a rate quoted without an interval tells you less than it appears to.

Because benchmarks are usually balanced and clean, while real inboxes are neither. Test sets hold roughly equal numbers of phishing and legitimate mail, and the phishing in them is a fixed historical collection. Raise the legitimate side to real-world proportions and the same detector's precision falls, because its false positives now outnumber its true ones.

They fail in opposite directions, so the accuracy question has no clean answer. Rules are precise about the pattern they were written for and blind to everything else. Models generalize to wording nobody wrote a rule for and cannot explain their own reasoning. Running both and combining the verdicts is the practical answer, not picking a winner.

Cross-corpus transfer is measuring a detector on a dataset it was not trained on. Scores typically drop sharply compared with a same-dataset test, which suggests part of what a model learns is the shape of its training corpus rather than the shape of phishing. A benchmark score from a same-dataset split is therefore optimistic by construction.

Ask four questions: which dataset, what was the class balance, what is the false-positive rate on legitimate mail, and does that rate come with a confidence interval. A claim that cannot answer all four is marketing copy rather than a measurement, however many decimal places it carries.

A confusion matrix is a two-by-two count of outcomes, and the false positives sit in the cell where the detector predicted phishing and the true label was legitimate. The cell diagonally opposite holds the true negatives, legitimate mail correctly cleared. Every rate is a ratio over these cells, so the false-positive rate is that one cell divided by the legitimate column, which is why it can be read without looking at the phishing side at all.

There is no universal threshold, because the cost lands on the user's attention rather than on a server. A rate is good if the person still reads the alerts after a month, which is a lower bar than it sounds: a single percent of a few thousand legitimate emails a week is dozens of false alarms. The independent comparison in this article measured two mature filters near 1% of legitimate mail and a third at 4.5%. Our own measured rate is 3.4%, with a 95% interval of 2.0% to 5.9%.

More from the blog