Skip to main content

IT risk scoring: qualitative, semi-quantitative and the ordinal-scale trap

A written-out 4 by 4 scoring scale, a worked example, and an honest account of why multiplying ordinal scores misleads and what to do about it.

Who this is for: anyone who has to put a number or a colour on an IT risk and then defend it to someone who was not in the room. That includes risk owners, security managers and the person preparing the figures for a board pack. You do not need a statistics background, but you will get more from this if you already keep a risk register.

Three ways to score, and what each one buys you

Before choosing a matrix, be clear about what kind of number you are producing.

ApproachWhat you assignWhat it is good forWhere it breaks
QualitativeLabels such as low, medium, highFast triage; workshops with non-specialistsLabels mean different things to different people; no arithmetic is valid
Semi-quantitativeOrdinal scores (for example 1 to 4) for likelihood and impact, usually combined into a bandA consistent ranking across many risks with modest effortScores are ranks, not measurements; combining them is lossy
QuantitativeA frequency or probability range and a loss range in moneyComparing a risk with the price of treating it; board-level money conversationsNeeds data or calibrated estimates; easy to produce false precision

Most organisations use the middle row for the register and reach for the third only for the handful of risks where the money decision is large. That is a sensible split, provided you know what the middle row cannot do. The rest of this guide is about using it carefully.

The scale used on this site, written out

The risk matrix tool and every example on this site use a 4 by 4 scale. A score is likelihood multiplied by impact. Four levels are deliberately few: an even number removes the safe “middle” option that makes three-level scales cluster, and it is easier to define four things precisely than seven.

Likelihood is judged over a stated horizon. Pick one (twelve months is common), write it at the top of the register, and use it for every entry.

LevelLabelDefinition (twelve-month horizon)
1UnlikelyWould need several independent controls to fail or an unusual combination of circumstances. No near miss in your own history.
2PossibleCould plausibly happen. Has happened to comparable organisations or you have had a near miss, and at least one effective control is in place.
3LikelyExpected to happen at least once in the horizon given the current controls, or the control is known to be weak or inconsistently applied.
4Almost certainHappening now, recurring, or the exposure is open and known with no effective control.

Impact is judged by the worst realistic consequence for the organisation, not the worst imaginable one. The anchors below are examples built on duration and reach so they work without inventing money figures; replace them with thresholds that mean something to your business, ideally expressed as a share of revenue, budget or margin.

LevelLabelExample anchors (replace with your own)
1MinorAbsorbed in normal operations. A workaround exists. No customer, regulator or contractual consequence.
2ModerateA service is degraded or unavailable for part of a working day. Internal rework. Limited exposure of internal data.
3MajorA critical service is down for days, or a contractual or regulatory obligation is breached, or personal data of customers is exposed. Visible to customers.
4SevereProlonged loss of a critical service, loss of data that cannot be recovered, or consequences that threaten a major contract, licence or the organisation’s viability.

The bands are fixed by the score:

ScoreBandReached by
1 to 3Low1×1, 1×2, 1×3, 2×1, 3×1
4 to 6Medium1×4, 2×2, 2×3, 3×2, 4×1
8 to 9High2×4, 3×3, 4×2
12 to 16Critical3×4, 4×3, 4×4

Ties are broken by impact first (the more severe consequence ranks higher), then by likelihood, then by order of entry. That rule is arbitrary, but it is stated and applied consistently, which is the property that matters.

A worked example

Five illustrative risks, scored with the scale above:

IDRisk (short form)LIScoreBandRank
R1Remote access without MFA4416Critical1
R2Admin accounts unreviewed for a year3412Critical2
R3Billing supplier outage248High3
R4Backups never restore-tested428High4
R5Laptop lost, disk unencrypted326Medium5

R3 and R4 both score 8. R3 ranks above R4 because its impact is higher (4 against 2), which is the stated tie-break. Notice what the tie-break is quietly saying: a rare, severe event outranks a frequent, modest one at the same product. You may disagree, and some organisations tie-break on likelihood instead. Either is defensible if it is written down before the scores are known. Choosing the rule after you see which risk you want on top is how a register loses its credibility.

Why multiplying ordinal scores misleads

An ordinal scale tells you order, not distance. Level 4 is worse than level 3, but it is not “one more unit” of anything, and it is certainly not twice level 2. Multiplying the scores treats them as if they were measured quantities. That has predictable consequences that you should know about before you present a matrix to anyone.

Different risks, same score. Under the scale above, a risk with likelihood 1 and impact 4 scores 4, and so does one with likelihood 2 and impact 2, and one with likelihood 4 and impact 1. These are very different: a once-in-a-generation event with severe consequences, a moderate event at moderate odds, and a constant small nuisance. The score places them in the same band. If you only ever read the band, you lose the shape of the risk. This is why the tool prints both factors and why the treatment suggestions look at which factor is driving the score.

Clusters and cliffs. Because only nine product values are possible with two four-point scales (1, 2, 3, 4, 6, 8, 9, 12, 16), risks bunch at a few scores, and the band boundaries create cliffs. A risk scoring 6 is Medium; one scoring 8 is High; yet the underlying difference may be one disputed judgement call about whether impact is 3 or 4. Be suspicious of any decision that hinges on a boundary.

Order can be wrong. Published critiques of risk matrices, including L. A. Cox Jr.’s 2008 paper in the journal *Risk Analysis* titled “What’s Wrong with Risk Matrices?”, argue that categorical matrices can mislead in several ways, including ranking a risk with a larger expected loss below one with a smaller expected loss. Douglas Hubbard’s book *The Failure of Risk Management* makes a related argument for practitioners. You should read the original sources rather than rely on this summary, but the practical lesson is robust: a matrix is a prioritisation aid for a conversation, not a measurement.

Hidden inputs. Two people can assign different likelihoods to the same risk because one was thinking “at least once ever” and the other “in the next year”. The product hides the disagreement. Always record the horizon and the reasoning.

Using the method responsibly

None of this means abandoning matrices. It means using them with their limits in view.

  1. Define the levels in writing, before scoring. The tables above are a template. If a definition could be read two ways, tighten it.
  2. Score likelihood and impact separately, then multiply. Never go straight to “this feels High”.
  3. Score the inherent risk and the residual risk separately. Inherent assumes no effective control; residual accounts for what is genuinely in place. Be honest about which controls you have *tested*.
  4. Score independently first, then compare. If three people each write a number before discussing, you find the disagreements that matter. If the most senior person speaks first, you find out what the most senior person thinks.
  5. Record the reasoning in a line. “Likelihood 3: restore never tested; one failed restore in 2024” is worth ten times a bare 3.
  6. Review the extremes. Look at every risk with impact 4 regardless of its score. Low-likelihood, high-impact items are exactly where products mislead.
  7. Re-score after treatment only if a control changed. Moving a number because it is uncomfortable is not treatment.

When to move to money

Move from scores to ranges of money when a decision is large enough to deserve it: choosing between treatments, sizing insurance, or answering “how much should we spend on this?” You do not need a model to start. Use ranges, not points.

A simple annualised form: expected annual loss is the loss per event multiplied by the number of events per year. Suppose your best honest estimates for a particular outage are a loss per event of 80,000 to 200,000 in your currency and a frequency of once in five to once in twenty years (0.05 to 0.2 events a year). The product spans roughly 4,000 to 40,000 a year. That is a wide range, and the width is the honest message: it tells you the estimate is not good enough to justify a very precise number. It still lets you compare against, say, an annual cost of 10,000 for a control that you believe would reduce the frequency by half.

Structured methods exist for this, including FAIR (Factor Analysis of Information Risk), which the FAIR Institute promotes; this site refers to it as a method and has no affiliation with that organisation. Whatever method you use, the same rule applies: show ranges, show assumptions, and say who made the estimate.

Common mistakes

  • Changing the scale between review cycles so trends cannot be compared.
  • Using the matrix to rank risks that are not comparable, such as a single system outage against a compliance gap.
  • Treating a green cell as “safe” rather than “currently not prioritised”.
  • Scoring the threat instead of the risk: “ransomware is a 5” says nothing about your exposure.
  • Scoring in a meeting dominated by one voice.
  • Presenting a heat map to a board with no stated scale. See board reporting for a format that avoids this.

What to do next

Copy the three tables above into your register template as the scoring key, with your own impact anchors. Score your current top ten entries separately for likelihood and impact, then paste them into the risk matrix tool to see the ranking and the plot; its output is a working estimate and not an audit or certification. Then take the top-ranked entries to risk treatment and decide what you will actually do about each.