Machine learning phishing detection means training a model on real phishing and legitimate emails so it learns the statistical shape of a scam, rather than matching fixed keywords. The model reads the subject, sender, and body, turns that text into numbers, and outputs a probability for each class it was trained on. A threshold converts that probability into a verdict.
- A phishing model learns patterns from labelled examples instead of matching rules someone wrote by hand.
- It reads text only: subject, sender, and the first 1,500 characters of the body, plus any rule-engine findings you pass it.
- The model's output is essentially one-hot, so the "confidence" percentage it shows is a decision, not a calibrated probability.
- Its headline accuracy usually measures a benchmark, not an inbox - our published 99.58% is scored on a test split that is 88.7% URL samples.
- Models and rules fail in opposite directions, which is the argument for running both and combining the verdicts.
What machine learning phishing detection actually is
Machine learning phishing detection is a classifier trained on labelled examples, not a set of rules somebody wrote by hand. The distinction sounds academic until you watch the two fail. A rule that flags "your account has been suspended" catches that exact phrase and nothing else. A model that has seen thousands of scams can flag a message using words nobody thought to write a rule for.
We built one of these models and ship it in our extension, so this article is written from the inside rather than from the literature. That matters here, because the numbers most pages quote about phishing detection turn out to measure something narrower than the claim they support - as the accuracy section below shows.
The stakes are not theoretical. The Anti-Phishing Working Group recorded 971,181 unique phishing attacks in Q1 2026 alone, up from 853,244 the previous quarter. No manual review process scales to that volume, and a blocklist only covers domains someone has already identified. Phishing also covers a much wider range than mass-mailed scams - our guide on what phishing is and how the main types differ breaks that range down.
So the question worth asking is not whether a model beats a rule. It is what the model is actually reading, and what it does with it.
What a phishing model actually reads
A phishing model reads text: the subject line, the sender's display name and address, and the opening of the body. Nothing else. It is not looking at attachments, it is not rendering the HTML, and it is not checking the sender domain against a registry. Everything it knows has to arrive as characters.
That constraint shapes the engineering more than any modelling decision. When you run Deep Scan in our extension, this is the block that reaches the model:
HEURISTIC FLAGS: - [high] sender-impersonation: display name claims "PayPal", domain is paypa1-secure.com - [medium] urgency-language: "within 24 hours"
Subject: Your account has been limited From: PayPal Service (service@paypa1-secure.com) Body: <first 1,500 characters of the visible message text>
Two details matter more than they look. The body is capped at 1,500 characters, and we measured what the cap costs: scoring the same samples with and without it moved recall by roughly 0.2 percentage points, 99.1% uncapped against 98.9% capped. Raising the limit is not a lever worth pulling.
The other detail is the flags block at the top. The model sees what the rule engine already found, so a suspicious sender domain it could never detect from text alone becomes part of its input. It is a second opinion, not a blind one.
One thing the model does not see is content hidden in the email with CSS. Since we split visible text from hidden text, the body it receives contains only what a reader would see. Hidden content reaches it indirectly, and only if a specific rule fired and reported it.
How an email becomes a verdict
Text becomes a verdict in three steps: tokenisation, attention, and a threshold. The email is split into tokens, each token becomes a vector of numbers, and a transformer weighs those vectors against each other to score every class the model was trained on.
Our model is a DistilBERT classifier with four output classes: legitimate email, phishing email, legitimate URL, and phishing URL. This is text classification in the plain sense: characters go in, a class comes out, and nothing about the sender's infrastructure is visible to it. It is a single-label model, meaning exactly one class wins per email. The model card's "multilabel" wording is wrong, and treating it as multilabel has already caused one silent, total failure of this feature.
The label names come from the dataset the model was trained on. We treat the dataset's labels as the reliable description when the two disagree, because they match the single-label behaviour we actually observe.
Then a decision threshold converts the output into a verdict. Worth being exact about what is thresholded: the model adds up the probability it assigns to its two phishing-labelled classes - phishing email plus phishing URL - and thresholds that total. Above 0.6 is scam, above 0.3 is suspicious, and 0.3 or below is safe.
Here is the part almost nobody tells you. That number is not a confidence. When we scored the shipped model against labelled emails, its per-class output turned out to be essentially one-hot: on 95% of samples the winning class's score sits above 0.99 with every other class near zero.
The model emits a decision, not a calibrated probability, and it fails calibration in the technical sense: the score does not track the real frequency of phishing among emails like the one being scored.
So when a panel shows "safe, 99% confident", that 99% does not mean the model weighed the evidence and settled at 99% certainty. It means the model picked that class, and it picks nearly everything that way. We tested whether a rule keyed on the margin between the top two scores could produce a useful "uncertain" verdict. It fired on almost nothing, and that measurement closed off the idea.
Machine learning vs heuristic rules for phishing detection
Rules and models fail in opposite directions, which is the real argument for running both rather than picking a winner. A rule is precise about the thing it was written for and blind to everything else; a model generalizes across wording it has never seen and cannot give you a faithful account of its own reasoning.
| Aspect | Heuristic rule | Machine learning model |
|---|---|---|
| What it can use as input | Anything the code can read: headers, links, domains, HTML structure, hidden text | Text only, as assembled before the call |
| Unseen wording | Misses it unless a rule happens to match | Generalizes from learned patterns |
| Explains its verdict | Yes - names the exact check that fired | No - reasoning is generated after the fact |
| Typical failure | False negatives on novel phrasing | False positives on legitimate mail that resembles a scam template |
| Cost per email | Instant, runs on every message | Seconds, and a model download on first use |
That last row is why our rule engine runs automatically on every email while Deep Scan is manual. It is not a quality ranking. The rules are cheap enough to run everywhere; a transformer inference is not, and running it on every message would trade a real cost for a marginal gain on mail the rules already handle.
What the accuracy number actually measures
A detection model's headline accuracy usually measures something narrower than the claim it gets used to support. Ours is a clean example. Our model's published score is 99.58%, and that is a score on its benchmark's test split - a split that is 88.7% URL samples. It describes URL classification, not email accuracy, and we do not quote it as an email figure.
| The number | What it actually measures | What it does not tell you |
|---|---|---|
| 99.58% | Accuracy on the benchmark's test split | Anything specific about emails - 88.7% of that split is URL samples |
| About 99% of phishing caught | Recall on the benchmark's two email classes | Whether the same holds on targeted attacks absent from the data |
| About 3% false positives | 3.4%, 95% confidence interval 2.0 to 5.9%, on the legitimate email class | How often it fires on your inbox - that class is not a sample of real mail |
| 0.2pp recall cost | Effect of the 1,500-character body cap, measured on the same samples | Nothing - this one is a genuine measurement, and it settles the question |
We measure this with a harness that scores the shipped model against labelled emails and reports every competing decision rule side by side, outside the extension bundle. It exists because a scoring bug once survived a fully green test suite. The tests agreed with the code, nothing compared either to reality, and every phishing email and every legitimate email came back "safe, 100% confident".
What we still cannot tell you is how the model performs on a real inbox. The benchmark's legitimate class is not a sample of real mail, and measuring against a live mailbox was declined as a matter of privacy. Treat the measured numbers as benchmark measurements, not as predictions of real-inbox performance.
There is a statistical reason for that caution, not just a methodological one. Precision - the share of a detector's alerts that turn out to be genuinely phishing - depends on how common phishing is in the stream being filtered, not only on how good the detector is. A false-positive rate measured where roughly half the samples are phishing does not transfer to an inbox where almost nothing is. Same model, same threshold, very different precision.
Why real email gets flagged as phishing
Legitimate email gets flagged because it genuinely resembles a scam template. A delivery notice, a bank statement, a forum notification, and the phishing email imitating them all share the same shape: an unfamiliar sender, a call to action, and an account reference. The detector has no way to know that this particular one is real.
This is the single most common question people bring to detection after they meet one. In r/phishing, a user posted a Samsung forum notification that had been flagged as a possible phishing attempt - an ordinary subscription reply, stamped with a warning by a different detector. The explanation in the replies was a mail-authentication mismatch on forwarded mail, not a content judgement at all.
That case is worth separating from ours, because two different things get called a false positive. Authentication checks compare SPF and DKIM records; a mismatch says the sending path looks wrong, and legitimate forwarded mail trips it constantly. A content model like ours never sees those records. It flags mail because the wording looks like a scam, which is why our false positives land on receipts, statements, and delivery notifications rather than on forwarded forum replies.
We probed this directly with a small set of hand-written samples. A CI notice, a shipping notice, a newsletter receipt, and a bank statement scored between 0.73 and 0.999, all far above the scam threshold.
Those specific numbers show where false positives concentrate, not how often they happen. The examples were written by someone who knew what a phishing template looks like, which selects for exactly the mail the model misreads. The 3.4% from the harness is the better estimate of frequency.
Where machine learning phishing detection fails
The uncomfortable part is that on transactional mail, no threshold separates real messages from phishing ones. They interleave. Adjusting the cutoff trades false positives for false negatives rather than removing either, which is why our results panel reports what the model observed instead of issuing a verdict.
False positives also arrive at the loudest setting available. In the harness run, all 12 misclassified legitimate emails landed on "Scam" and none on "Suspicious", so the failure mode is never hedged. Roughly 1 legitimate email in 20 to 50, on a benchmark whose legitimate class is not a sample of real mail.
The second blind spot is targeted impersonation. Business email compromise often contains no malicious link and no attachment at all, just a convincing request from someone who appears to be a colleague. There is very little for a text classifier trained on consumer phishing to latch onto, and very little for a rule to match either.
If you want the human-readable version of those tells, our guide on how to spot a phishing email covers the red flags directly, including the ones no automated system checks. CISA's guidance on recognizing and reporting phishing takes the same position from the other direction: reporting is part of the defence precisely because automated detection has limits that people are the last line against.
Model drift: why a detection model goes stale
Model drift is the widening gap between the examples a model was trained on and the messages it meets in production. The model does not decay on its own - the same input always produces the same output. What moves is everything around it: attackers change their templates, legitimate senders change their formatting, and the mix of mail in a real inbox shifts with the seasons.
The general name for that in machine learning is distribution shift; drift is what it looks like from the defender's side. The failure is quiet. A drifted model does not start erroring; it starts being wrong in the same confident, one-hot way it was always wrong. There is no signal in its output that says "I have not seen anything like this before", which is exactly why the one-hot finding from earlier matters more than it first appears.
Our model is pinned to a specific version and is not updated periodically. That is a deliberate trade-off rather than neglect. Pinning means a scan behaves the same way tomorrow as it did today, which matters for a security tool people build habits around. It also means the model ages, and we track that as an open question rather than pretending it is settled.
This is the structural argument for keeping a rule engine alongside the model. Rules age too, but a rule's age is visible: it is a line of code someone can read, date, and replace. A model's age is distributed across 67 MB of weights, and the only way to see it is to measure against fresh labelled data on a schedule.
How we combine the AI verdict with the rule engine
We never let the model's verdict stand alone. Four fixed rules decide which signal wins, and they are deliberately asymmetric: the model can raise a floor but cannot lower one. Any critical rule finding, or a rule-engine score at or above the scam threshold, produces a scam verdict regardless of the model.
A high-severity finding plus a model verdict of safe forces the result to at least suspicious, so the model cannot clear an email the rules flagged strongly. A model verdict of scam against a rule verdict of safe upgrades only to suspicious, which is what stops the model creating a scam verdict on its own. In every other case, the rule engine's verdict stands unchanged.
Disagreements surface in both directions rather than being silently resolved. When the model clears an email the rules flagged, you see a note explaining the conflict and recommending you trust the rules. When the model flags an email nothing else objected to, you see a note saying that finding is a hint rather than a verdict. Neither system gets a blind veto.
| Rule engine | Model | Final verdict |
|---|---|---|
| A critical finding, or a score at or above the scam threshold | Anything, including safe | Scam |
| A high-severity finding, below the scam threshold | Safe | Suspicious |
| Safe | Scam | Suspicious |
| Anything else | Anything | Whatever the rules decided |
If you want the architecture underneath that - why the model runs in your browser rather than on a server, and what that costs and buys - our companion piece on on-device versus cloud AI phishing detection covers that decision in full. This article is about the model itself; that one is about where it runs.
Final verdict - phishing detection machine learning
Machine learning genuinely adds something rules cannot: the ability to recognize a scam written in words nobody anticipated. The useful version of that claim is much narrower than most marketing makes it. The model reads text and only text, its confidence score is a decision rather than a probability, its headline accuracy usually measures a benchmark rather than an inbox, and its false positives land on ordinary transactional mail.
The honest summary is that neither approach is the answer on its own. Rules are cheap, explainable, and blind to novelty. Models generalize, and cannot explain themselves or tell you when they have gone stale.
Running both, and defining precisely which one wins a disagreement, gets you the coverage of one and the accountability of the other. If you want to see what that looks like in practice, our write-up on how Email Scam Checker works walks through the rule engine, the model, and the fusion rule between them.