Detecting sophisticated phishing: from the mirage of perfect scores to an honest benchmark
In short: on public corpora our phishing detector showed an F1 of 0.99. A leakage audit revealed that a quarter of the test set was in fact memorised during training, the score was a mirage. Evaluated on a held-out benchmark of hard cases (BEC, mobile money, quishing, FR+EN), a baseline model dropped to near chance and blocked 37% of legitimate mail. By rebuilding the data, diversity and hard legitimate cases rather than volume, performance on this internal benchmark reaches an AUC around 0.98. We present that figure for what it is: a controlled internal benchmark, not real-world production performance.
The mirage of perfect scores
Public phishing datasets are old, mostly English, and riddled with near-identical emails: the same template resent hundreds of times. A random split scatters these near-twins across both sides, training and test, so the model recognises at test time what it memorised during training. The resulting F1, close to 1, does not measure the ability to generalise to unseen phishing.
So we audited leakage before anything else. The result: 24% of our test set had a near-twin in training, up to 79% on some sources. The 0.99 was not performance, it was memorisation. It is the same methodological discomfort as in network detection: a figure that looks too good almost always hides a protocol that is too lenient.
An honest judge: the hard-case benchmark
To measure what really matters, we built a benchmark of difficult cases, written by hand and kept strictly out of training: CEO fraud with no link or attachment, payment-redirection fraud, mobile-money phishing (Orange Money, Wave), QR-code quishing, in both French and English, alongside deliberately tricky legitimate mail (genuine internal payment requests, security notices, statements).
The verdict was harsh and healthy. On these realistic cases, the baseline sat at chance level (AUC ≈ 0.67) and blocked more than one legitimate email in three. Detection excelled at yesterday's phishing, yet remained blind to sophisticated phishing, precisely what AI tools now produce.
Sovereign data, not millions of clones
The temptation is to 'generate millions of examples'. We rejected it: a million clones reproduces the mirage at scale. The real lever is diversity. We built a compositional, sovereign, offline generator that assembles each message from varied phrasings, local contexts (mobile money, banks, tax administration) and realistic addresses, half phishing, half hard legitimate mail, deduplicated, and filtered so it never overlaps the evaluation benchmark. It is the hard legitimate cases that drive false positives down.
| Metric (internal benchmark) | Baseline | Iteration 3 | Iteration 4 |
|---|---|---|---|
| AUC, hard cases | 0.67 | 0.98 | 0.98 |
| Recall, phishing caught | 0.61 | 0.87 | 0.91 |
| False positives, legit blocked | 37% | 5% | 10% |
| Recall, English | n/a | 0.63 | 0.88 |
What these numbers say, and do not say
We refuse to turn this benchmark into a marketing claim. Training relies on synthetic data and evaluation on a set we wrote ourselves: the figure is therefore likely optimistic. A 10% false-positive rate is still too high to automatically block business mail; in production this signal must first warn, not silently delete. The only proof that counts will come from real mail.
Next: real data
Synthetic data let us bootstrap and measure honestly, it has a ceiling. The next step is real data, consented and anonymised at the source, and an annotation loop where the analyst corrects verdicts. Two hundred labelled real emails are worth more than two hundred thousand extra synthetic ones. That is how a good prototype becomes a product worthy of trust.