loading

Whether it worked

What departments say their own rules will do.

  1. 01the forecastyou are here
  2. 02the verdict on itRPC opinions
  3. 03the review, years onpost-implementation reviews
  4. 04the independent auditNAO reports

Before making a rule, a department publishes what it thinks the rule will cost and achieve, marked against a standard six-line grid. This is every assessment this year — and how many carry no grid at all.

We read all 131 impact assessments published in 2026 and coded every scorecard we found. Only 43 of the 91 that should carry one do. What those 43 say, and don't say, is below; every rating is in the heatmap. This is stage one of Whether it worked: the forecast, written before anything happens. The Regulatory Policy Committee's verdict on it is stage two — that page is the marker's verdict, this one is the homework.

43 of 91

Fewer than half of the forward-looking impact assessments published in 2026 carry a scorecard at all. The thing this page is about is mostly missing. 48 assessments that should have one don't — and nothing on the public record flags which.

131 entries listed for 2026 40 post-implementation reviews, which don't need one 91 assessments that do 48 without
43 scorecards 12 departments 12 also ruled on by the RPC Last updated · checked weekly

The pattern

Six things the 2026 scorecards show.

Most assessments still don't have one

131 impact assessments were published in 2026. Forty are post-implementation reviews, which use a different template and aren't expected to carry a scorecard. That leaves 91 forward-looking assessments that should — and 48 of them don't. The flagship transparency device of the reformed framework is absent from more than half the documents it was designed for.

The gap isn't random, and it isn't evenly spread. Roughly a third of the missing ones are Northern Ireland assessments, which run on a separate devolved template that predates the scorecard. The rest are Great Britain assessments that simply don't have one: some state explicitly that they sit outside the Better Regulation Framework, some are tax impact notes, and a long tail are older summary sheets that still print an NPSV and an EANDCB but no directional ratings at all. Nothing in the public listing distinguishes any of these from an assessment that does carry a scorecard — you have to open all 131 PDFs to find out, which is what we did.

It matters because the scorecard is the only place a department has to state, in public and in one word, which way each impact points. Where it's missing, that statement was never made.

Nothing is bad overall

Twenty-four of the 29 distinct assessments rate their own effect on total welfare as positive. Three say neutral, one says uncertain, one leaves the line blank. Not a single department, across a full year, published a scorecard saying its measure would make the country worse off. A scale on which the top box is ticked five times in six is not measuring much.

The whole cost lands in one column

Business is the only line where negatives appear at all — eleven of twenty-nine. Nine of those eleven sit directly beside a positive welfare rating. That pairing is the framework doing its job: it makes the transfer visible, and it says out loud that the case for these measures rests on gains landing somewhere other than the regulated firm.

Part B is mostly blank space

Eleven of the twenty-nine say nothing but neutral, uncertain or nothing at all across all three wider-priority lines. International trade is neutral in seventeen of twenty-nine. Where Part B does carry a signal it is almost entirely the decarbonisation line. The business environment and international questions are being answered, but they are not being used.

The same assessment is published several times

Fourteen of the 43 rows are repeat publications — one assessment attached to two or more statutory instruments. The higher education tuition fee assessment appears four times; the lifelong learning fee limits assessment three. Anyone counting published impact assessments without deduplicating is over-counting by about a third.

Departments aren't filling it in the same way

Part B has its own vocabulary — supports, may work for, may work against — and several departments answer it with Part A's positive/negative wording instead. One assessment merges two Part B lines into one; another drops the business and household rows entirely. The template is standard; the practice isn't yet.

Every rating

Positive on the left, empty on the right.

One row per impact assessment, one cell per scorecard line. Part A says who is affected; Part B says whether the measure helps or hinders three wider government priorities. Hover any cell for the department's own wording and the money behind it.

Regulatory scorecards, 2026

Part A rates the direction of impact on three groups — Positive, Neutral, Negative or Uncertain. Part B asks whether the measure supports or works against three wider priorities, on a five-point scale from Supports to May work against. A blank cell means the line was left unanswered or could not be read from the published PDF.

Regulatory scorecards published in 2026, scored part by part.
Part A · stakeholder impacts Part B · wider government priorities
Impact assessment Total welfare Business Households Business
environment
International Natural capital
& decarb.
Part A ++ Positive = Neutral −− Negative ? Uncertain Not stated
Part B ++ Supports + May work for = Neutral May work against ? Uncertain Not stated
Flags ▣ published more than once ◆ non-standard scorecard / the RPC's verdict, where it has published one

Marked homework

What the RPC made of the same documents.

Twelve of these assessments have also been through the Regulatory Policy Committee, so for those we have both halves: the department's own statement of which way the impacts point, and an independent verdict on whether the analysis behind it stands up. Our companion page, what the RPC keeps finding, covers the verdicts in full.

Reading them side by side, the striking thing is how little the two have to do with each other.

A failed assessment still has a cheerful scorecard

The Future Homes and Buildings Standards assessment was rated not fit for purpose: the RPC red-rated both the identification of options and the justification for the preferred way forward. Its scorecard says the overall impact on total welfare is positive. Nothing about the red verdict is visible on the scorecard, because the scorecard doesn't ask about the quality of the evidence — only about the direction of the answer.

The RPC signs off on scorecards that all say the same thing

It rated the scorecard section Satisfactory in ten of the twelve, Good in one and Weak in one. So on the RPC's own measure these are, with one exception, adequate scorecards — and they are also almost uniformly positive. "Satisfactory" is a judgement about whether the box was filled in properly, not about whether the answer is informative.

The one rated Good is the one that admits a cost

The single scorecard the RPC rated Good is the National Minimum Wage uprating — which is also one of the few to record a negative impact on business, with an EANDCB of £157.6m and a business net present value of −£1.1bn attached to it. The scorecard that scores best is the one carrying the biggest number with a minus sign in front of it.

The weakest column isn't on the scorecard at all

Monitoring and evaluation is the section the RPC criticises most often, across its whole caseload. In this matched set it rated the planning appeals assessment Very Weak on it. A reader of that assessment's scorecard would never know: the scorecard has six lines, and none of them asks whether anyone plans to find out if the measure worked.

The twelve matched assessments

Sorted worst-first on the RPC's judgement. The left pair of columns is the RPC's, transcribed from its opinion; the six on the right are the department's own, transcribed from its scorecard. Hover any cell for the wording behind it.

RPC Fit for purpose Not fit for purpose G Good S Satisfactory W Weak ↺ green only after an initial failure

A note on what this is and isn't. Twelve pairs is not a sample you can run a correlation on, and we haven't tried. What the matched set shows is structural rather than statistical: the scorecard and the RPC opinion are measuring different things, and a department can satisfy one while failing the other. Pairs were matched by hand on department, title and date, and only included where all three agree — several near-misses were left out rather than guessed at.