Email Privacy Statistics: How to Read the 2026 Data
Where credible email privacy and breach statistics come from, how each source is measured, why the widely quoted numbers disagree, and how to use the data to make decisions about your own inbox.
Email privacy articles are saturated with statistics, and a surprising share of them cannot be traced to anything. A number appears in a vendor blog post, gets quoted without its year, gets rounded in the next retelling, and eventually circulates as common knowledge with no surviving link to a study.
This guide takes the opposite approach. Rather than reprinting figures, it maps the primary sources that actually measure email-related privacy and security incidents, explains what each one counts and does not count, and shows why two credible reports can produce very different headline numbers about the same year.
Deliberately, this article does not restate specific percentages. Breach and phishing figures change annually and are frequently misquoted; the responsible thing is to send you to the current edition of each source rather than freeze a number here that will be wrong in six months. Every source below publishes free, dated, methodology-documented editions.
Why Most Quoted Email Statistics Are Unreliable
Most circulating statistics fail one of four tests: they have no named source, no year, no stated sample, or no definition of the thing being counted. A figure missing any of these cannot be verified or compared, and is usually a rounded retelling of an older number.
The failure is structural rather than dishonest. Vendors summarise a report, bloggers summarise the vendor, and each step drops a qualifier. 'Of confirmed breaches analysed in this dataset' becomes 'of breaches', which becomes 'of companies'. By the fourth retelling the claim describes a population nobody measured.
Definitions are the sharpest problem in email specifically. 'Phishing attack' can mean an email sent, an email delivered, an email clicked, or an incident with confirmed loss — four numbers that differ by orders of magnitude and are routinely presented interchangeably.
A quick test before trusting any figure: can you name the organisation that measured it, the year, the population sampled, and the definition used? If any of the four is missing, treat the number as illustrative at best.
Key takeaways
- Demand source, year, sample and definition before trusting a figure.
- 'Phishing attacks' can mean four different things; check which.
- Rounded numbers without links are usually third-hand retellings.
The Primary Sources Worth Reading
Five sources do original measurement on email-related privacy incidents: the Verizon Data Breach Investigations Report, IBM's Cost of a Data Breach Report, the FBI's Internet Crime Report, the UK ICO and Australian OAIC breach registers, and Have I Been Pwned's breach index. Each is free and dated.
These five are worth knowing individually because they answer different questions. The DBIR characterises how incidents happen. IBM prices what they cost. IC3 counts what victims report to law enforcement. The regulators publish what organisations were legally obliged to disclose. HIBP records which specific credential dumps became public.
None of them is a general population survey, and that is the crucial caveat. Every one samples a specific slice — investigated incidents, surveyed organisations, self-reported crimes, mandatory filings, published dumps — so their numbers are not interchangeable even when they describe the same phenomenon.
| Source | Counts | Main blind spot |
|---|---|---|
| Verizon DBIR | Investigated incidents and confirmed breaches, categorised by pattern | Skews to organisations that engage investigators |
| IBM Cost of a Data Breach | Surveyed cost per breached organisation | Self-reported costs; large-enterprise weighting |
| FBI IC3 | Crimes reported to the FBI by victims | Massive under-reporting; US-centric |
| ICO / OAIC registers | Breaches organisations were legally required to notify | Only covers notifiable breaches in those jurisdictions |
| Have I Been Pwned | Credential dumps that became publicly available | Silent breaches never appear |
Why Credible Reports Disagree
Reputable reports disagree because they use different denominators. One counts incidents investigated, another counts organisations surveyed, a third counts victims who filed a report. Each denominator excludes a different population, so the same underlying reality yields very different percentages.
Consider the simplest question: how common are email-driven breaches? Measured against investigated incidents, email-based intrusion looks dominant, because those are the incidents serious enough to investigate. Measured against all organisations, most were never breached at all in a given year, and the same phenomenon looks rare.
Under-reporting compounds the gap. Law-enforcement figures capture only victims who filed a complaint, which is a small and non-random fraction. Regulator registers capture only breaches meeting a notification threshold. Neither is wrong; both are partial.
The practical response is to read directionally rather than precisely. Where several independent sources with different methodologies point the same way, the direction is trustworthy even when no single number is.
- Different denominators produce different percentages from identical facts.
- Self-reported data systematically undercounts.
- Agreement across methodologies is stronger evidence than any single figure.
What the Sources Consistently Agree On
Across every major source and every recent year, three findings hold: email is the dominant delivery channel for initial compromise, stolen or reused credentials are a leading cause of breach, and the human element is involved in a large majority of incidents. These directions are stable even as the exact figures move.
Email's role as the delivery channel is the most robust finding in the field. Whether the vector is described as phishing, business email compromise or pretexting, the message arrives in an inbox, and the address that received it had to exist in some dataset first.
Credential reuse is the second constant. Once an address and password pair appears in one dump, automated credential-stuffing attempts follow across unrelated services. This is why a breach at a forum you forgot about can become a problem at a service that was never breached at all.
The third constant is timing: detection and containment take substantially longer than most people assume, which means a leaked address is typically in circulation well before anyone is notified. That lag, more than any single percentage, is the reason to limit how many databases hold your primary address.
Key takeaways
- Email is the delivery channel for most initial compromises.
- Credential reuse turns one breach into many.
- Detection lag means notification arrives long after exposure.
Reading a Statistic Properly: A Worked Method
To evaluate any email statistic, locate the primary source, check the publication year, identify the sample population, read the definition of the counted event, and confirm whether the figure is a rate or a raw count. If a step is impossible, the statistic cannot support a decision.
Start at the end of the chain. Follow the link in the article you are reading; if it leads to another article rather than a report, keep going. Most chains terminate within two or three hops at a named report — or at nothing, which is itself the answer.
Then check the edition. Annual reports are superseded every year, and a figure from an edition three years old describes a threat landscape that has since changed considerably. Publication year should always appear alongside a quoted number.
Finally, ask whether the figure would change your behaviour if it were half as large, or twice as large. If not, the number was never the decision driver — the direction was. That question filters out most statistics-driven anxiety.
| Step | Question | Fails if |
|---|---|---|
| 1 | Who measured it? | No named organisation |
| 2 | When? | No year, or an edition superseded |
| 3 | Who was sampled? | Population undefined |
| 4 | What counted as an event? | Definition unstated or shifting |
| 5 | Rate or count? | Presented ambiguously |
Turning the Data Into Decisions
The consistent findings support three actions: use unique credentials with a password manager, enable phishing-resistant multi-factor authentication on accounts that matter, and give a disposable or aliased address to any service you have not decided to trust long term.
None of these depends on a precise statistic. They follow from the direction all the sources agree on: addresses leak, leaked pairs get reused, and email is where the attempt arrives. Reducing the number of databases holding your primary address reduces the surface for all three.
Compartmentalisation also gives you diagnostic value. When a per-service address starts receiving unrelated mail, you learn exactly which company leaked or sold it — information no public report can give you about your own data.
Monitoring closes the loop. Checking your primary address against a public breach index periodically tells you when a rotation is warranted, and it is the one place where a statistic about you specifically is both accurate and actionable.
- Unique passwords per service, stored in a manager.
- Phishing-resistant MFA on email, finance and identity accounts.
- Disposable inboxes for trials and downloads, aliases for durable accounts.
- Periodic breach checks on your primary address.
Frequently Asked Questions
Why does this article not quote specific percentages?
Because email privacy figures are revised annually and are the most commonly misquoted numbers in the field. Freezing a percentage here would guarantee it becomes wrong and then gets recirculated without its year. The sources listed publish current, dated editions with documented methodology — read the figure at the source.
What is the most reliable source for breach statistics?
There is no single best source. Verizon's DBIR is the strongest for how breaches happen, IBM's report for what they cost, and national regulator registers for legally notified breaches in their jurisdictions. Use them together and trust the direction they agree on rather than any one headline number.
How do I check whether my own email address has been breached?
Search your address in a reputable public breach index such as Have I Been Pwned. It only covers dumps that became public, so a clean result is not proof of safety, but a positive result is a concrete signal to change that password everywhere it was reused and enable multi-factor authentication.
Are phishing statistics comparable between reports?
Usually not. Reports variously count emails sent, emails delivered, users who clicked, and incidents with confirmed loss. Comparing a click-rate from one report with an incident-rate from another produces a meaningless conclusion, which is how many viral statistics originate.
Does a bigger reported breach number mean things are getting worse?
Not necessarily. Mandatory notification laws expanded substantially over the last decade, so more breaches are disclosed than before even where the underlying rate is flat. Rising counts partly measure improved reporting, which is why methodology notes matter more than year-over-year headlines.
What single change reduces my exposure the most?
Not reusing passwords, enforced by a password manager. After that, limiting how many services hold your primary address — disposable inboxes for anything transient and aliases for durable accounts — shrinks the number of databases that can leak an identifier tied to you.
Sources & further reading
Related Reading
Explore the blogPut It Into Practice
After reading the strategy, the fastest next step is to test the workflow with a real disposable inbox. That makes the comparison practical instead of theoretical and helps you see whether the verification flow, delivery speed, and privacy tradeoffs fit your use case.