Skip to main content

How to run a phishing test that teaches instead of punishing

Seven steps, five mistakes worth avoiding, and the two numbers actually worth measuring.

A phishing test — Hook calls it a simulation, because a test implies someone passes or fails it — sends a safe version of a real attack to your own people. Done well it rehearses a reporting reflex that shortens real incidents. Done badly it teaches people that admitting a mistake is expensive, and you lose the reporting you were trying to build. The difference is almost entirely in the details below.

The two numbers that matter

Most programs report a click rate and stop. Click rate is easy to drive to zero by sending easier scenarios, which is why it flatters programs that are not working.

Reporting rate

What share of people who received the message reported it. This is the behavior you want, and unlike click rate it cannot be inflated by making the scenario easier — an easier message is easier to report too.

Time to report

How long from delivery to the first report. This maps directly onto what an incident costs: an attack reported in four minutes is a containable event, and the same attack reported in four hours usually is not.

Track click rate too — just never alone, and never attached to a name. See report rate and time to recognition for how Hook defines them.

Seven steps to run your first phishing test

  1. 01

    Decide what the test is for before you send anything

    Write down the one question you want answered. "Do people report suspicious email, and how fast?" is a usable question. "What percentage of our staff is gullible?" is not — it has no action attached to either answer.

    The question determines everything downstream: which scenario you pick, what you measure, and what you do with the result. Teams that skip this step end up with a click rate, no idea whether it is good, and nothing to do on Monday.

  2. 02

    Tell people the program exists — just not when

    Announce that simulations are part of how the organization trains, before the first one goes out. Do not announce the dates. Surprise about timing is the realistic part; surprise about the existence of the program is what breeds resentment.

    This is the single biggest predictor of whether a program survives its first year. An unannounced first simulation reads as a trap, and people who feel trapped stop reporting — which destroys the metric that actually matters. A one-paragraph note from a leader is enough.

  3. 03

    Get your reporting path working first

    Before the first simulation, confirm there is a one-click way to report a suspicious message and that someone is on the other end of it. Test it yourself with a real message.

    If reporting is a three-step process involving a forwarded attachment and a ticket queue, your reporting rate measures the friction of your process, not the alertness of your people. Fix the path first or you will misread the result.

  4. 04

    Start with an easy, plausible scenario

    The first simulation should be a common pattern — a shared document, a delivery notice, a password expiry — sent to everyone, not a targeted impersonation of your CEO.

    A hard first simulation produces a high click rate, a demoralized workforce and a number you cannot compare to anything. Start where most real attacks start. Difficulty is something you earn the right to increase once reporting is healthy.

  5. 05

    Make the moment after a click teach something

    Someone who clicks should immediately see a short, private explanation of what the message was and which tell gave it away. Not a score. Not a notification to their manager.

    This is the only part of a simulation that changes behavior. The click itself teaches nothing; the thirty seconds afterwards teach everything. Keep it private, keep it under a minute, and keep blame out of it entirely.

  6. 06

    Measure reporting, not just clicking

    Record how many people reported the message and how quickly, alongside how many clicked. Reporting rate and time-to-report are the numbers that predict how expensive a real incident will be.

    Click rate alone is a vanity metric that can be driven to zero by sending easy simulations. Reporting rate cannot be gamed the same way, and it maps directly onto incident cost: an attack reported in four minutes is a different event from one reported in four hours.

  7. 07

    Repeat on a cadence and change one variable at a time

    Run monthly rather than annually, and vary the tactic — urgency, authority, helpfulness — rather than escalating difficulty every round.

    Recognition decays. USENIX SOUPS research on reinforcement found the effect of a single training moment fades within months, which is why an annual test measures memory rather than habit. Varying tactic instead of difficulty tells you which manipulation your team is still vulnerable to, which is actionable in a way that one blended number is not.

Five mistakes that cost more than they teach

Each of these produces a worse number next quarter, not a better one.

Sending a fake bonus, layoff or payroll message

These produce the highest click rates and the worst outcomes. People who feel manipulated about their pay or their job security do not become more careful — they become less willing to trust security email at all, which is the opposite of the goal. Several organizations have made national news doing this.

Publishing a leaderboard of who clicked

Naming people converts a learning moment into a public penalty. The measurable result is fewer reports, later, because the cost of admitting a mistake now exceeds the cost of staying quiet.

Treating a click as a disciplinary event

Once a click can affect a review, people hide them. Hidden incidents are the expensive ones. Discipline belongs to deliberate policy violation, not to falling for a message that was engineered by professionals to be fallen for.

Running one test a year and calling it a program

An annual simulation tells you how people performed on one morning. It does not build a habit, and it is usually scheduled near an audit, which means it measures preparation rather than behavior.

Scoring individuals and following them around with it

Attaching a persistent number to a person reads as a character judgement, and people manage the number rather than the risk. Measure resilience and vulnerability by tactic across a team instead — that tells you what to train next, which a per-person score never does.

What to send first

Pick a pattern people meet weekly rather than one that would fool a security team. Our ten phishing email examples break down what each pattern pretends to be, the tactic underneath it, and the tell that gives it away — any of the first few make a reasonable opening scenario.

If you would rather not build and schedule this yourself, Hook Security runs the simulations for you: scenarios go out on a cadence you set, a private training moment fires the instant someone clicks, and the reporting numbers come back as a report you can hand to an insurer or an auditor.

Common questions about phishing tests

Decide the question you want answered, announce that the program exists without announcing dates, confirm your reporting path works, then send a common, plausible scenario to everyone. Show anyone who clicks a short private explanation immediately, and record reporting rate and time-to-report alongside click rate. Repeat monthly, varying the tactic rather than escalating difficulty.

In the United States, running simulations against your own workforce on company systems is standard practice and is expected by most cyber insurers and auditors. Rules differ elsewhere — works councils in parts of the EU, and employee-monitoring and data-protection law generally, can require notice or consultation first. Treat this as a question for your own counsel rather than as settled.

Published benchmarks put an untrained baseline in the high teens to low thirties percent, falling into single digits with a sustained program. Treat those numbers cautiously: click rate depends almost entirely on how hard the scenario was, so it is comparable to your own previous results and to nothing else. Reporting rate is the more honest measure of whether a program is working.

Monthly. Recognition decays measurably within months, so an annual test measures memory rather than habit. A monthly cadence with a rotating tactic keeps the reflex rehearsed and produces a trend you can actually read.

They should immediately see a brief, private explanation of what the message was and the tell that gives that pattern away, and nothing else should happen. No manager notification, no score, no leaderboard. The teaching moment is the entire value of the exercise, and it only works if admitting the click is safe.

Tell them the program exists; do not tell them the schedule. Announcing the program up front is what separates a training exercise from a trap, and it protects the reporting behavior you are trying to build. Announcing the dates would make the result meaningless.

They describe the same exercise. "Test" is the common search term; Hook uses "simulation" because a test implies a pass or fail applied to a person, and grading people is what makes them stop reporting. The word matters less than whether your program treats a click as a data point or as a verdict.

See what your reporting rate actually is.

Thirty minutes, a live account, and a straight answer about where your people stand today.