Short Disclaimer or Copyright notice.
Feature Article: Improving Your Forecasting Skills: Defining Questions, Establishing Base Rates, and Weighing New Evidence to Update Probabilities


Over the years, I’ve learned a lot about forecasting the behavior of complex socio-technical systems (e.g., organizations, industries, financial markets, nation states, and global systems) from a wide range of very smart people, including Pierre Wack, Philip Tetlock, Doyne Farmer, Marvin Cohen, Gary Klein, and many others. Clearly I am a fox, not a hedgehog!

To help our subscribers become better forecasters, this month’s feature article focuses on three issues, the importance of which is often underappreciated. These are (1) How to define forecasting questions, (2) How to establish a prior or baseline forecast, and (3) How to weigh new qualitative evidence and update your prior.

Defining Your Forecasting Questions

Many of the questions about the future that we encounter are broadly defined – for example, “How is the US-China relationship going to evolve?” or, “Will I outlive my retirement savings?”

Answering these questions requires that we break them down into sub-questions that are more amenable to the application of forecasting methods.

Developing these sub-questions requires us to have or develop some type of causal model that includes the key elements in a situation and how they interact with each other to produce an answer to the broad question we are trying to predict.

Effective sub-questions have three other characteristics: (1) they are phrased in a way that allows the use of probabilities to answer them; (2) they have a specific time horizon; and (3) they have clear criteria for establishing the accuracy of a forecast at or before the end of that time horizon.

For example, to forecast the relative returns on different asset classes, we focus on four sub-questions

(1) What is the probability that, 12 months from now, we will be in the High Uncertainty Regime (defined as a decline in the FTSE All-World equity index, expressed in US Dollars, of at least 20% over the previous 12 months)?

(2) What is the probability that, 12 months from now, we will be in the High Inflation Regime (defined as an increase of 5% or more in the US Consumer Price Index over the previous 12 months)?

(3) What is the probability that, 12 months from now, we will be in the Persistent Deflation Regime (defined as a decline in the US Consumer Price Index over the previous 12 months)?

(4) What is the probability that, 12 months from now, we will be in the Normal Regime (defined as not being in any of the other regimes)?

In our causal model, these macro regimes are the primary driver of asset class returns. In turn, these macro regimes result from the interplay of a series of higher-level drivers that operate in a rough chronological sequence (with multiple feedback loops), including changes in technological, economic, national security, social, and political conditions.

It is important to note that most socio-technical systems are both complex and adaptive. Forecasting the future outcomes produced by such systems is extremely difficult, not only because they typically have many cause-effect relationships, but also because many of these relationships are time-delayed and/or non-linear. Moreover, the adaptive actions of agents (both human and algorithmic) within such systems often cause the structure of the system itself to evolve over time.

In modeling terms, the sources of uncertainty include randomness, the parameter values for different variables and relationships, and the structure of the model itself.

As Robert Hoffman and Gary Klein have shown (in their series of papers on “Explaining Explanation”) the challenge posed by the complexity of socio-technical systems also extends to our attempts to explain their past behavior, which is often a key source of the mental models we use to forecast their future outcomes.

To be sure, advancing technology is slowly making this task easier (e.g., agent based modeling and simulation and a variety of artificial intelligence methods). But using models to forecast the future behavior of complex socio-technical systems still has a very long way to go (e.g., “Priority Challenges for Social and Behavioral Research and Modeling”, by Davis et al from RAND).

At this point, when trying to anticipate the behavior of complex adaptive socio-technical systems the best we can hope for are forecasts that are “coarse-grained”, but accurate, particularly as the time horizon extends. Our goal is to produce accurate forecasts of broad macro regimes, rather than very specific events that could contribute to the development of a given regime.

Establishing Your Prior/Baseline Forecast

Having defined your forecasting questions, the next challenge is developing your prior (to use the Bayesian term) or initial/baseline forecast. In many cases, this turns out to be more difficult than it first seems.

In the easiest case, a baseline forecast can be constructed using easily obtained data on the frequency of comparable historical events – e.g., the percentage of times a US equity market index fell by 20% or more over rolling 12 month periods since a given starting date.

When direct evidence like this is unavailable, you must use other techniques. One approach I have found very useful starts with the search for analogies to the forecasting problem at hand and determining if there is better historical data available for them.

If there is not, I ask myself what would be the likely shape of the distribution of outcomes if they were available. This is a very important step, because our baseline mental model of what a distribution looks like – the normal (Gaussian, Bell Curve) distribution we learned in statistics class – often does not accurately describe the distribution of results produced by complex social systems (as Nassim Taleb reminded us in Fooled by Randomness and The Black Swan).

Instead, outcomes produced by systems in which social learning or influence occur are best described by a power law function or Pareto-type distribution, with a large number of small outcomes and a few very large ones (e.g., see, Boisot and McKelvey, “Extreme Events, Power Laws, and Adaptation” and Andriani and McKelvey, “From Gaussian to Paretian Thinking”).

However, the histories of complex social systems are also filled with plausible counterfactuals, because the outcomes we observe are the result of interacting situational forces, human decisions, and randomness (or, if you will, luck). For this reason the use of analogical reasoning to establish a prior must be balanced with an analysis of the forces at work in the case at hand, and the forecast probabilities they imply. As a final step, the probability produced by analogy must be combined with the probability derived from analysis of the current situation. As a general rule, the more similar the case is to the analogy used, the greater the weight that should be put on the latter.

A third way to establish a prior is to ask yourself what the conventional wisdom is on the question you are trying to forecast. This approach can be particularly useful in situations of high uncertainty, where, as John Maynard Keynes noted, commonly accepted “conventions” are used when no better information is available. Evidence that is subsequently gathered can then be used to test the accuracy of the conventional forecast, as, for example, described by Rappaport and Mauboussin in their 2001 book, Expectations Investing.

Weighing New Evidence and Updating Your Prior Probabilities

Having established your baseline/prior forecast probability estimate, the next challenge is determining how you should adjust it after receiving new information.

In my experience on the Good Judgment Project, I found that this process of evidence collection, evaluation, and weighting was critical to the accuracy of my forecasts.

In a world of information overload, the first challenge is deciding which information to target for collection. I found that doing a pre-mortem on my baseline forecast was extremely useful in this regard. A pre-mortem uses a technique called “prospective hindsight” to highlight weaknesses in our forecasts. Research has shown that when we attempt to explain the past, we get to a much more specific level of detail than we do when trying to forecast the future. Hence, the pre-mortem technique asks you to imagine that at some time in the future, your forecast has been shown to be wrong. It then asks you to write down why this occurred – the important information or dynamics you missed, and what you could have done differently to avoid being wrong. Having used it both individually and in groups, I can assure you it is a very powerful technique (as more systematic research has also found).

As a general rule of thumb, the greater the number of assumptions you use in your argument, and the more uncertain they (and perhaps the forecast logic itself) are, the lower should be the resulting forecast probability.

Doing a pre-mortem on our forecast highlights the assumptions and/or logic in your forecast about which you are most uncertain, and which would therefore most benefit from collecting additional information.

There are three questions that must be asked about each new piece of information you obtain, before it is combined with others and weighed to determine by how much you should adjust your most recent forecast probability.

The first question is whether the information is relevant to the forecasting question at hand. If your collection effort is targeted, it should be.

The second question is whether the information is credible. In a world of fake news and echo chambers, this has become a very non-trivial issue. While intentionally deceptive information has always been a concern for professional intelligence analysts, it is now a concern for everyone.

Assuming new information is relevant and credible, the third step is to weigh its value and impact on your prior forecast probabilities.

Broadly, there are three systematic approaches to weighing evidence (which itself is an ambiguous phrase, with no agreed upon meaning).

In the 17th century, Sir Francis Bacon posited that the weight of evidence for or against a hypothesis depends on both how much relevant and credible evidence you have, and on how complete your evidence is with respect to those matters that you believe are critical for evaluating a hypothesis.

Bacon recognized that we can be “out on an evidential limb” if we draw conclusions about the probability a hypothesis is true based on our existing evidence without also taking into account the number relevant questions that are still not answered by the evidence in our possession. We typically fill in these gaps with assumptions, about which we have varying degrees of uncertainty. In this context, the value of a new piece of information depends on the degree to which it either reduces the number of assumptions you use, or increases your confidence that they are accurate.

In the 18th century, Reverend Thomas Bayes invented a quantitative method for using new information to update a prior estimated probability (degree of belief) in the truth of a hypothesis.

”Bayes Theorem” says that given new evidence (E), the updated (posterior) probability that, in light of this new evidence, a hypothesis is true p(H|E) is a function of the conditional probability of observing the evidence given the hypothesis p(E|H), times the prior probability that the hypothesis is true p(H), divided by the probability of observing the new evidence p(E).

The “Likelihood Ratio” is a critical Bayesian concept. It is equal to the probability of observing a piece of evidence if a hypothesis is true, (E|H), divided by the probability of observing the evidence if the hypothesis is false, p(E|not-H). The greater the Likelihood Ratio for a piece of new evidence, the greater is its information value, and thus the larger should be the difference between your prior and posterior probabilities. When it comes to evaluating new evidence, the Likelihood Ratio is a very valuable heuristic that is easy to intuitively apply.

In the 20th century, Arthur Dempster and Glenn Shafer developed a new theory of evidence weighing.

Assume a set of competing hypotheses. For each of these hypotheses, a new piece of evidence is assigned to one of three categories: (1) It supports the hypothesis; (2) It disconfirms the hypothesis (i.e., it supports “Not-H”); or (3) it neither supports nor disconfirms the hypothesis. To relate this back to Bayes, in cases (1) and (2), the evidence has a high Likelihood Ratio; in the case (3) it does not.

The accumulated and categorized evidence can then be used to calculate a lower bound on the belief that each hypothesis is true (based on the number of pieces of supporting evidence and their quality), as well as an upper bound (equal to one less the probability that the hypothesis is false, again, based on the evidence that disconfirms the hypothesis, and its quality). This upper bound is also known at the plausibility of each hypothesis.

The difference between the upper (plausibility) and lower (belief) probabilities for each hypothesis is the degree of uncertainty associated with it. Hypotheses can then be ranked based on their degrees of uncertainty. Similar to the Likelihood Ratio, in the Dempster-Shafer context the value of a new piece of information is proportional to the change it produces in the uncertainty associated with one or more hypotheses.

While there are quantitative methods for applying all of these theories, they can also be applied qualitatively, to quickly produce an initial conclusion about which of a given set of hypotheses is most likely to be true, and by how much you should adjust a forecast probability.

For example, in our analytical process, we use the same probability categories as the US Intelligence Community:

Almost Certain: 95% or more
Very Likely: 80% - 95%
Likely: 55% - 80%
Even Chance: 45% - 55%
Unlikely: 20% - 45%
Very Unlikely: 5% - 20%
Almost No Chance: 5% or less


These categories provide a starting point for using evidence to formulate and later update estimated forecast probabilities. The probability adjustments should be proportional to the relevance, credibility, and information value of new evidence that is received.

However, one of the key lessons from the Good Judgment Project was that the best forecasters make probability distinctions that are finer than these seven broad categories (see, “Small Steps to Prediction Accuracy” by Atanasov et al).

Another lesson was that individual forecasters’ probability estimates were often too close to the 50% “toss-up” category that is the least useful to policymakers. In other words, individuals typically did not give enough weight to the information they used to make their forecast. The GJP team compensated for this (and increased forecast accuracy) by combining individual forecasts and then making the resulting probability more extreme. Subscribers can download this “extremizing” equation from our website.

The final – and critical – step in using new evidence to update your prior probability (to a so-called “posterior”) is to repeat the pre-mortem process.

This has two important benefits in addition to focusing your subsequent information collection activities. First, it forces you to make explicit the logic, evidence, and assumptions that underlie your forecast. Second, by forcing you to recognize the greatest uncertainties in your forecast, it reduces your overconfidence in its accuracy.

Caveat #1: The Dangers of Social Information

Broadly speaking, there are two kinds of social information you can receive. One will help improve your forecast accuracy, and the other very likely will not, particularly in highly uncertain situations.

The first is information from other people preparing forecasts on the same issue you are, which details their logic, evidence, assumptions, and conclusions, and/or challenges your own. The Good Judgment Project showed that this team-based approach can improves forecast accuracy.

The second type of social information simply tells you which is the most popular forecast. This is likely to worsen the accuracy of your forecast, especially in highly uncertain situations where accurate forecasts are most valuable. The reason for this lies deep in our evolutionary past, when increased uncertainty raised our anxiety about being cast out from our group, and hence made us much more likely to conform to its prevailing opinion. This herding process can lead a large number of people to accept a forecast that was originally made on the basis of very little information, or using weak logic and assumptions.

Caveat #2: The Critical Role of Surprise

Surprise is fleeting feeling that has underappreciated importance for updating forecast probabilities. As Daniel Kahneman explained in his book, “Thinking Fast and Slow”, the experience of feeling surprised is transitory because human beings naturally try to minimize energy and time intensive cognitive effort.

Surprise is triggered by our perception of information that either is at an extreme end of the range of outcomes that our mental models of the world predict (such as very rapid or large change), or by an observation that is inconsistent with the models themselves (e.g. awakening to a green sky). To minimize cognitive effort and reduce uncertainty, we subconsciously try to adjust our mental models to enable surprising information to cohere with our existing beliefs. Most of the time, this process occurs automatically; it is only when coherence can’t quickly be reestablished that we become conscious of feeling surprised.

When we are trying to accurately forecast a future outcome, and particularly when the distribution of that outcome follows a power law, it is easy to see how our natural suppression of surprise can get us into trouble.

For this reason, when startled by surprise, you should immediately try to write down what triggered it, before the feeling disappears. This enables you to later more carefully consider the implications of that trigger.


If you have any questions about anything we have written in this issue, please don’t hesitate to get in touch, at contact@indexinvestor.com