Crisis RadarUpdated 16 Aug · next cycle 17 AugRequest access
Briefings ·

Our first test against reality

A month ago we published 560 predictions and promised to mark them against what actually happened, in public, whatever the result. Here are the marks.

Blair Cowan · Crisis Radar · 4 min readcalibration · methodology · track-record · forecast-resolution
The short of it

In June we published 560 predictions across 35 countries. In July we checked them against what actually happened. On questions about year-level conditions, such as whether a country's inflation stays above crisis levels, we did slightly better than someone who simply quoted history. On the strict 30-day questions, nothing on our list happened during the month, so the only thing that month could measure was our false alarms. We are publishing all of it, including the parts that flatter no one, and some marks will change as slower official data arrives. Every change will be logged here.

On 22 June this system published 560 predictions. For each of 35 countries it answered 16 questions: will there be a coup attempt? Will a new war start? Will inflation pass 40 percent for the year? Will the government impose new restrictions on moving money out? Each answer was a probability with a date attached, and we made one promise about them: when the date arrives, we check every prediction against what actually happened and publish the result, whatever it says.

The first batch came due on 22 July. This is the result.

What actually happened

Of the 560 predictions, 27 of the events came true. 358 did not. The remaining 175 cannot be checked yet: they concern slow-moving measures of governance, built on academic datasets that only publish once a year, so we set them aside entirely rather than quietly counting them as correct.

How we mark ourselves

The test is simple to state. For every prediction, compare us against the most honest lazy alternative there is: ignore the news completely and just quote how often this event has happened historically in similar countries. If our predictions are not closer to reality than that, we are adding nothing, and the product has no reason to exist.

Making that comparison fair turned out to require one important repair. Our questions come in two kinds. Some ask whether something will happen inside a 30-day window: a coup, a new conflict, new sanctions. Others really ask whether the year will cross a line: inflation above 40 percent, the economy shrinking more than 5 percent. Marking the year-questions against 30-day history made us look better than we are, so we now mark the two kinds separately, each against its own fair yardstick, and we publish no single combined headline. The combined number was the flattering one, and it is gone.

The marks

On the year-questions, we beat history, modestly. Thirteen of those events came true, and our predictions were closer to what happened than the historical odds were. Most of that edge came from recognising conditions that were already under way, such as a country already deep in an inflation crisis, and saying so plainly rather than deferring to long-run averages that said such things are rare. For readers who track the formal scoring: our error score was 0.062 against the yardstick's 0.072, where lower is better.

On the 30-day questions, we cannot claim anything yet. Not one event on that list happened during the month, in any of the 35 countries. When nothing happens, simply quoting history is almost impossible to beat, because history said "almost certainly not" and was right every time. The only thing that month could measure was our false alarms, the probability we spent on things that never occurred. We spent some. So the honest statement is: on the strict 30-day questions, this month proves nothing about our skill, and we will not pretend otherwise.

Two of the sixteen questions, new sanctions and new capital controls, have no historical yardstick built yet, so they sit outside the test entirely. Their predictions are recorded and checked like all the others, but there is nothing fair to compare them against until that reference data exists. We say this on the scorecard rather than folding them in quietly.

These marks belong to the old version of the system

The June predictions were made by the system as it stood in June. The process of checking them is precisely what exposed its worst habits: it treated countries already in crisis as if the crisis were unlikely, it confused ongoing wars with new ones, and it was overconfident at the top of its range and timid in the middle. We have rebuilt those parts since. The system now checks what is already true in a country before it predicts, applies the strict definitions of war and onset correctly, runs three independent models instead of one, and passes every result through a correction curve fitted to nearly 500 checked predictions. None of that improves this month's marks, and it should not: these are the old system's marks. The rebuilt system's predictions start being checked from the end of August.

Some marks will change, and you will see it happen

Three recession predictions were marked using IMF projections because the real figures do not exist yet; they will be re-marked in October when they do. Twenty are waiting on academic conflict data that publishes next year. When any mark changes, the score changes, and every change will be logged here with the reason and the before-and-after numbers. A scoreboard that moves silently is worthless. One that corrects itself in public is the point.

The next batch of predictions comes due on 21 August, and the largest one in late September. Both will be marked the same way and published here, whatever they say.

For live country coverage, forecasts, and the daily briefing, request access →