Back

Calibration and Collective Judgment

How superforecasters make calibrated predictions, how crowds aggregate uncertainty, and how meta-predictions reveal expertise hidden inside a consensus.

Nebula

Superforecasting is the practice of making unusually accurate probability estimates about uncertain future events. It is not clairvoyance and it is not a single model. It is a system: define questions that can resolve, start from base rates, decompose the problem, update as evidence changes, keep score, and combine judgements without hiding disagreement. A related technique, meta-prediction, asks forecasters what they expect other forecasters to believe. That second-order signal can reveal information that a simple crowd average misses.

Bottom line

Good forecasting is an evidence-and-feedback discipline. Superforecasters are partly selected and partly trained; crowds become useful only when their information remains sufficiently diverse; and meta-predictions can help identify who knows something the crowd does not. Every claim must ultimately survive a scoring rule and a resolved outcome.

What superforecasting means

A forecast should state a probability for a clearly defined event by a specific time. “The company is likely to struggle” is an opinion. “There is a 35% probability that the company reports year-on-year revenue decline in its next audited results” is a forecast that can be resolved and scored. The precision is not a claim to certainty; it makes uncertainty explicit.

A behavioural model estimates what a particular actor may do in a particular context; our research on behavioural simulation examines how those actor-level estimates should be grounded and evaluated. Superforecasting instead asks how a forecaster repeatedly forms and improves probabilities across resolvable questions.

The modern evidence base comes largely from geopolitical forecasting tournaments. In one study, five university teams recruited forecasters, elicited probability estimates, and aggregated them across hundreds of questions. Mellers and colleagues found that training, teaming, and tracking improved both calibration and resolution. Training taught probabilistic reasoning and the use of reference classes. Teams exchanged evidence and challenged rationales. Tracking identified consistently strong performers and placed them in enriched environments.

Follow-up research described the best performers as partly discovered and partly created. They combined cognitive ability and open-mindedness with task-specific skill, motivation, frequent updating, and effective collaboration. The label “superforecaster” therefore describes measured performance within a forecasting system—not a permanent title awarded for confidence, seniority, or one spectacular call.

Calibration, resolution, and scoring

Forecast quality has at least two dimensions. Calibration asks whether events assigned 70% probability occur about seven times in ten over a suitable set of questions.Resolution asks whether the forecasts meaningfully separate likely outcomes from unlikely ones. Predicting 50% on every binary question may look cautious, but it contains little information. Predicting 99% repeatedly may look decisive, but one surprise is heavily penalised.

Proper scoring rules reward honest probabilities. The Brier score for a binary event is the squared distance between the forecast probability and the outcome, conventionally coded as zero or one. Lower is better. Log scores penalise assigning extremely low probability to events that occur even more sharply. A tournament should specify its scoring and resolution rules before forecasters see the outcome.

Scores need context. Question difficulty, forecast timing, and opportunity to update all affect performance. Comparing one person's early forecast with another person's final forecast is not meaningful. A robust ledger preserves every timestamped probability, the information available at the time, and the final resolution source. This is the forecasting equivalent of the append-only traces described in our research on simulating society.

How superforecasters reason

Start with the outside view

The outside view asks how often comparable events occurred before considering the vivid details of the current case. If the question concerns whether a merger will close, start with completion rates for comparable deals. If it concerns a product deadline, begin with the organisation's historical delivery rate. The reference class anchors the forecast before a compelling narrative pulls it toward an extreme.

Decompose the question

Large questions often hide several smaller ones. A forecast about an election may depend on turnout, candidate eligibility, economic conditions, coalition behaviour, and polling error. A market forecast may depend on attention, positioning, a catalyst, and the price response. Breaking the problem apart exposes assumptions and makes it easier to identify the evidence that would move the final probability. Our article on prediction markets explains how contracts turn similarly explicit outcomes into prices.

Update in proportion to evidence

A new fact should change the probability only as much as its diagnostic value warrants. Repeating the same report across five outlets is not five independent pieces of evidence. A forecaster should record why the probability moved, which assumption changed, and what evidence would reverse the update. Small updates are often appropriate; refusing to update is not a sign of conviction when the information set has changed.

Search for disconfirmation

Strong forecasters treat their estimate as a working model. They look for base rates that contradict their first intuition, distinguish facts from interpretations, and ask what the opposing forecast explains better. The goal is not performative balance. It is to reduce the chance that one coherent story monopolises the probability space.

From individual forecasts to crowd forecasts

A crowd forecast is not automatically wise. Aggregation works when people contribute partially independent information and the mechanism combines it sensibly. If everyone reads the same source, copies the visible consensus, or shares the same model architecture, the apparent crowd may contain only one underlying signal repeated many times.

The simplest aggregate is a mean or median probability. More sophisticated systems can weight forecasters by historical calibration, question-domain performance, or contribution to prior crowd forecasts. Some methods extremise an average—moving it away from 50%—because individual estimates can be conservative. But extremising a biased or highly correlated crowd can amplify error. The adjustment should be validated out of sample rather than adopted because sharper probabilities look more informative.

Dynamic aggregation also matters. Forecasters update at different times, so a current aggregate should not treat a month-old estimate as equal to one made after a decisive event. A hierarchical model developed by Satopää and colleagues combined sparse, asynchronous updates from more than 2,300 forecasters across 166 international events and produced aggregates designed to remain sharp and calibrated in and out of sample. See their dynamic probability-aggregation research.

Meta-prediction: forecasting the forecasters

A first-order forecast asks, “What is the probability of the event?” A meta-prediction asks, “What probability will other forecasters assign?” The gap between the two can contain information. A person who expects the crowd to say 30% but independently says 60% may possess private evidence, misunderstand the problem, or simply be contrarian. Across many participants, meta-predictions help distinguish these possibilities.

Research on identifying experts when past performance is unknown uses each person's own probability and their estimate of the crowd's average. The proposed meta-probability weighting method performed better than comparison aggregators across 500 binary decision problems in the reported experiments. The useful signal was not confidence alone; it was the relationship between the forecast and the forecast of others' forecasts.

Later work on recalibrating aggregate probabilities with meta-beliefs uses meta-predictions to estimate a shared prior and adjust the crowd away from that prior rather than mechanically away from 50%. This addresses an important limitation of simple extremisation: the neutral midpoint is not necessarily the correct reference point.

The terminology needs care. In judgemental forecasting, meta-prediction usually means a prediction about other people's predictions. In machine-learning research,meta-forecasting can instead mean learning across many time series so a system can select or adapt a forecasting method for a new series. Both operate one level above an ordinary forecast, but they solve different problems. This article uses meta-forecasting only as the broad family name and meta-prediction for the second-order belief itself.

Can AI systems superforecast?

Language models can search, summarise, generate rationales, and produce probabilities, but fluent reasoning does not guarantee calibration. An AI system also risks knowledge-cutoff leakage, duplicated sources, fabricated evidence, unstable probabilities, and correlated errors across agents that share a base model.

A 2024 study on language-model forecasting built a retrieval-augmented system that searched for evidence, generated forecasts, and aggregated predictions. On questions published after the models' knowledge cut-offs, the system approached the aggregate performance of competitive human forecasters and surpassed it in some settings. The important result is architectural: retrieval, question filtering, multiple forecasts, and aggregation mattered.

AI forecasting should therefore be evaluated like any other forecasting system. Freeze the information boundary, preserve sources, record model and prompt versions, prohibit access to future resolutions, score every probability, and compare against simple baselines and human aggregates. Model diversity should be measured rather than assumed. Twelve agents prompted differently may still behave like one correlated forecaster.

Building a forecasting system

  1. Write resolvable questions. Define the event, deadline, data source, edge cases, and cancellation rules before collecting predictions.
  2. Elicit probabilities and rationales. Ask for a number, the strongest evidence, a reference class, assumptions, and the next observation likely to move the forecast.
  3. Collect meta-predictions selectively. When expertise histories are sparse or shared information may bias the crowd, ask what the forecaster expects others to predict.
  4. Aggregate transparently. Publish the raw median or mean alongside any weighted, recency-adjusted, or extremised result. Back-test the transformation.
  5. Update without rewriting history. Preserve every timestamped forecast and explain material changes rather than displaying only the latest number.
  6. Resolve and review. Apply the predeclared resolution criteria, compute proper scores, inspect calibration, and run postmortems on both successes and failures.

This loop is useful well beyond geopolitics. It can discipline product planning, risk management, market research, policy analysis, and scenario selection. Forecasts should inform decisions rather than impersonate them: the probability of an event, the value of acting, and the organisation's tolerance for downside are separate quantities. Compare that discipline with the practical tools in our stock sentiment stack, or browse the full research library.

Where forecasting systems fail

Vague questions cannot resolve cleanly. Correlated crowds create false confidence. Outcome leakage makes back-tests meaningless. Survivor selection celebrates a few dramatic winners without counting all failed calls.Persuasive rationales can be mistaken for accurate ones. Recent research found that eloquence and confidence increased persuasion without reliably indicating forecasting accuracy.

Forecasts can also change the systems they describe. A public probability may affect behaviour, attract attention, or alter incentives. Market sentiment is especially reflexive: a forecast about a narrative can become part of the narrative. Our work on measuring market sentiment separates public tone from positioning and fundamentals for precisely this reason.

Finally, no method removes irreducible uncertainty. Rare structural breaks, adversarial actors, missing data, and genuinely novel events can defeat good process. The purpose of disciplined forecasting is not to eliminate surprise. It is to make beliefs explicit, improve them through evidence, and learn when reality disagrees.

Sources and further reading

Research standard

Ask a question that can resolve. State a probability. Record the evidence boundary. Update when the evidence changes. Preserve the full history. Aggregate without pretending the crowd is independent. Then score the result. Forecasting becomes a science only when being wrong leaves a trace.

Written by

Hidden Systems Research

Social and behavioural systems research team. Hidden Systems studies how computational agents, behavioural evidence, networks, and uncertainty can be used to investigate complex social systems.

Research the systems behind collective behaviour

Explore more Hidden Systems research on social intelligence, behavioural modelling, and the forces shaping markets.