Red, Amber, Green Is Not a Number
Someone on the board asks what the exposure is worth, and the heat map cannot answer. It was never built to. Here is what goes wrong when ordinal scales get treated as arithmetic, what quantifying a risk actually involves, and which risks are not worth quantifying at all.

The slide is up. Five by five grid, the usual colours, eleven dots scattered across it with two sitting in the top right corner in red.
A board member asks how much the red one is worth.
There is a pause, because the honest answer is that the grid does not contain that information and never did. What comes out instead is a sentence about significant potential impact to operations and reputation. The board member nods, the meeting moves on, and nothing that happens in the next twelve months is different as a result of that slide having existed.
This happens in a lot of organisations, and the heat map usually gets blamed for it. That is slightly unfair. The heat map is a useful triage tool being asked to do a job it was never designed for.
The arithmetic does not work
Here is the technical problem, and it matters more than it sounds.
A five point impact scale is ordinal. It tells you that 4 is worse than 3. It does not tell you by how much, and the gaps are not equal. In most scales built in practice, the distance between 4 and 5 is enormous compared to the distance between 1 and 2, because the top of the scale is anchored to something existential and the bottom is anchored to an annoying week.
Now watch what almost every risk register does with those numbers.
It multiplies them. Impact times likelihood. Then it averages them across a portfolio, sorts by the result, and presents the top ten.
None of those operations are valid on ordinal data. A risk scoring 4 by 4 is not four times worse than one scoring 2 by 2. The average of a 5 and a 1 is not a 3 in any meaningful sense: it describes a risk that does not exist. And the sort order at the top of the list is often decided by where somebody drew a category boundary, not by anything about the risks.
You can do this arithmetic. Excel will not stop you. It just does not mean anything, and the number that comes out has an authority it has not earned.
What the board is actually asking
When someone asks what the exposure is worth, they are not being difficult and they are not asking for false precision. They are asking a capital allocation question, and it usually has this shape:
You want to spend two million on this control programme. What is the thing we avoid by spending it, and is it bigger than two million?
A colour cannot participate in that conversation. A range in currency can.
The answer does not have to be precise. "Somewhere between four hundred thousand and eleven million, most likely around one point eight" is an enormously more useful sentence than "high", because it can be compared to the cost of the control. That comparison is the entire point.
What quantifying actually involves
Quantification has an unfortunate reputation for being a modelling exercise that requires a statistician. Some of it does. The basic version does not.
Stop asking for one number. Ask for a range and a confidence. Not "what does this cost" but "what is the low end, what is the high end, and how sure are you". People are surprisingly well calibrated on ranges and terrible at point estimates, because a point estimate feels like a commitment and a range feels like an honest answer.
Separate frequency from severity. How often does this happen, and when it happens, how bad is it? These are different questions with different evidence behind them, and jamming them into one score destroys both. A high frequency, low severity risk and a low frequency, catastrophic one can land on the same square of a heat map while requiring completely different responses.
Model the distribution, not the average. The average outcome of a risk is often a number that will never actually occur. What matters for most decisions is the tail: not the typical bad year, but the bad year that would genuinely hurt. A distribution shows you that. A single figure hides it.
Record what actually happened. This is the step that gets skipped, and it is the one that makes everything else improve over time.
Loss events: the part everyone skips
When a risk materialises, there is an incident, and the incident gets managed. Systems come back, a report gets written, people move on.
What usually does not happen is anyone going back to the risk register entry that described this exact scenario and writing down what it actually cost.
That is a shame, because it is the only feedback your estimates will ever get.
Without recorded loss events, your risk assessments are opinions that never get marked. Estimates do not improve, they just get repeated, and the annual review becomes a ritual where last year's numbers get copied forward with small adjustments nobody can justify.
With them, you slowly acquire something valuable: evidence about how wrong your organisation tends to be, and in which direction. Most teams discover they are systematically optimistic about frequency and systematically pessimistic about severity. Knowing that about yourself is worth more than another scoring workshop.
Appetite, if it is going to mean anything
Most risk appetite statements are prose. "The organisation has a low appetite for risks affecting customer data."
Nobody can breach that. It contains no threshold, so no state of the world contradicts it.
An appetite that does anything has a number attached and a consequence when the number is crossed. Tolerance is the point at which something specific happens: an escalation, a decision, a named person being told. Not a paragraph, a trigger.
The same is true of indicators. A key risk indicator with no threshold is a metric on a dashboard that somebody looks at during the quarterly review, by which point the interesting movement happened months ago. A KRI with a threshold is a tripwire. It either fires or it does not, and when it fires the platform tells someone rather than waiting for them to notice.
When not to quantify
This is the part that usually gets left out of articles like this one, so: quantification is not the right answer for everything, and pretending otherwise is how programmes lose credibility.
Novel risks with no base rate. If something has genuinely never happened to anyone in a comparable position, a distribution built from nothing is theatre. Scenario analysis is more honest.
Risks where the real question is not financial. A regulatory breach with a personal liability attached to a named executive is not usefully expressed as an expected value. Neither is a safety risk. The decision is not being made on expected cost and modelling it that way is a category error.
The long tail. If you have four hundred risks in a register, quantifying all of them is a year of work that produces a spreadsheet nobody reads. Quantify the fifteen that drive decisions. Leave the rest qualitative and revisit them annually.
The heat map keeps a job here, and it is a good one: triage. It sorts the four hundred into the fifteen that deserve real analysis. That is genuinely useful. It is just not a board reporting tool, and the trouble starts when it gets promoted into one.
What this looks like in Sentinel Unity
The mechanics, stated concretely:
- Risk matrices and scoring formulas are configurable. A five by five is one option, not the assumption. If your organisation has a considered reason for a different shape, the platform holds it.
- Quantitative analysis with distributions produces a range rather than a point, so the tail is visible instead of averaged away.
- Loss events are recorded against the risk they belong to, which is what eventually lets you compare what you predicted to what occurred.
- Appetite and tolerance carry thresholds, so a breach is a state the platform can detect rather than a judgement someone makes later.
- Indicators carry their own thresholds and raise when crossed.
- Taxonomy is versioned, with a diff, because when the scale changes, everyone who compares this year to last year needs to know what moved and why.
That last one is less exciting than the others and causes more arguments than all of them combined.
Where to start
Take the three risks your leadership actually argues about. Not the top three by score, the three that generate real disagreement in the room.
For each, get a range instead of a rating. Ask two people separately for a low, a high, and how confident they are. Compare the two answers.
The gap between what two informed people believe about the same risk is usually the most interesting number in the entire register, and no heat map will ever show it to you.
Sentinel Unity supports configurable risk matrices, quantitative analysis with distributions, recorded loss events, and appetite thresholds that raise when breached. Request a demo, or read more about enterprise risk management.
See the platform on your frameworks
Request a walkthrough with our team, tailored to your entity structure and regulatory scope.
Request a demo