82,361 Forecasts, 3 Archetypes: Why the Hedger Wins in 2026
82,361 Forecasts, 3 Archetypes: Why the Hedger Wins in 2026
| Takeaway | Detail |
|---|---|
| Polymarket prices were 2026's best-calibrated public forecast | Poly Syncer's May 17, 2026 audit of every market resolving between January 2024 and May 2026 found just 2.1 points of mean absolute calibration error across all probability buckets, sampled 24 hours before resolution — ahead of sportsbook lines and polling averages. |
| The 50% fence-sitter loses even to a base-rate quoter | Answering 50% on every question delivers perfect calibration with zero discrimination — 'useless, but honest,' per EdgeMarket — so a lazy forecaster who merely quotes the historical base rate, say 11.2%, outscores the fashionable 0.5-hedger. |
| Stated 90% confidence is systematically inflated | Events people rate 90% certain typically occur only about 70% to 80% of the time (PredictEngine), and Convexly notes overconfidence clusters in the high-confidence bins of reliability diagrams. |
| Extremize crowds, never yourself — the Hedger's bounded edge wins | Extremization is a crowd correction, so individuals applying it simply import overconfidence; the winning Hedger instead anchors aggressively on base rates, bounds conviction (a genuine edge reads 55%, not 90%), and updates in log-odds — exploiting soft spots like systematically underconfident crypto-price binaries. |
Just 2.1 points of mean absolute calibration error separated Polymarket prices from reality across every probability bucket in Poly Syncer's May 17, 2026 audit — the sharpest publicly observable forecast tested, ahead of sportsbook lines and polling averages. But calibration, EdgeMarket warns, is 'the cheap half of the problem': a curve hugging the diagonal says nothing about what a forecaster should do with the next fresh question.
Here is the trap beneath the praise: the most expensive five words in forecasting are 'it could go either way.' The fashionable 0.5-hedger answers 50% on every question, pinning a Brier score at a guaranteed ceiling of mediocrity by pure arithmetic — perfectly calibrated, zero discrimination, 'useless, but honest.' Even a lazy forecaster who merely quotes the historical base rate — 11.2%, say — beats the fence-sitter, because discrimination cannot be repaired after the fact.
The loud 0.9-hedgehog fails from the other side: events people rate 90% certain typically happen only about 70% to 80% of the time, PredictEngine reports, and overconfidence piles up in the high-confidence bins of every reliability diagram. Worse, the trendy counsel to 'extremize everything' was derived for crowds, never individuals. What survives both collapses is the Hedger — the boring fox with a spreadsheet whose edge is aggressive base-rate anchoring, bounded confidence, and log-odds updates.
The 0.25 Trap
Glenn W. Brier defined the score in the Monthly Weather Review: subtract the outcome (1 if the event happens, 0 if it does not) from your stated probability, square the difference, average across questions. According to PredictEngine's primer, 0 is perfection, 0.25 is no skill—the exact score of guessing 50 percent throughout—and anything lower marks genuine forecasting ability. Now spring the trap. Answer 0.50 on every 2026 question and your penalty is (0.50 − 1)² = 0.25 when the event lands and (0.50 − 0)² = 0.25 when it misses. You post 0.25 whether you sweep the sheet or go winless. The safest-looking strategy locks in mediocrity before a single question resolves.
Allan Murphy's partition in the Journal of Applied Meteorology splits that score into reliability minus resolution plus uncertainty, and the middle term is where hedgers die. Reliability asks whether events you call at p actually occur near frequency p—EmergentMind formalizes it as E[(p − π(p))²]—while resolution asks how much your answers separate the outcomes that happen from the ones that don't. A flat 0.50 has zero spread across questions, hence zero resolution: it cannot tell the recession year from the normal one even holding the full answer sheet. Convexly's decomposition work shows two forecasters posting identical Brier totals—one clustered at the base rate, one trading some reliability error for real separating power. Only the second carries a signal anyone can score.
The quadratic penalty explains why. Worked examples Numa published on June 10, 2026 show a 0.9 call that lands scores (0.90 − 1)² = 0.01, while the same call that misses costs (0.90 − 0)² = 0.81. Loss scales with the square of your distance from truth, so the scoring rule quietly installs Aristotle's mean: expected loss is minimized by stating your evidence-weighted probability, never the center of the scale. The mean is not an ethical compromise here—it is the argmin.
The dominance proof follows in one line. Quote the historical frequency p as your answer and your expected Brier is p(1 − p): at p = 0.8 that is 0.16, against the hedger's locked-in 0.25. Since p(1 − p) reaches its maximum of 0.25 only at p = 0.5, committing to the mean of the evidence beats committing to the middle of the scale for every base rate except the exact coin flip. The folk advice—hover near 50/50 until certainty arrives—is therefore instructions for purchasing the maximum available penalty.
The deeper cost is informational. Resolution requires variance across your answers, so a forecaster whose twenty 2026 estimates all sit between 0.45 and 0.55 has declared zero knowledge: modesty-as-default deletes exactly the spread that makes someone measurable, and what cannot be measured cannot be improved. EdgeMarket said it plainly on August 15, 2026—the pure base-rate quoter is perfectly calibrated with zero discrimination, "useless, but honest."
Hence the fixed contest the rest of this guide argues over. The 2026 Calibration Test specifies: at least 20 resolvable yes/no questions spanning recession timing, CPI trajectory, AI-benchmark milestones, and midterm outcomes; every probability logged at a stated date; every answer scored by Brier at resolution. One sheet, one metric, no renegotiating the scoreboard midyear.
| Stated probability | Score if YES | Score if NO | Expected score* |
|---|---|---|---|
| 0.50 (chronic hedge) | 0.25 | 0.25 | 0.25 |
| 0.80 (base rate) | 0.04 | 0.64 | 0.16 |
| 0.90 (when true rate is 0.90) | 0.01 | 0.81 | 0.09 |
*Expected values assume outcomes arrive at the stated probability. Read down the column: evidence-anchored commitment at 0.16 beats the hedge at 0.25, and calibrated conviction at 0.09 beats both—which is why the test holds answers inside the 0.1–0.9 band rather than banning confidence. Build the sheet this week: twenty dated questions, quarterly Brier scoring, and the arithmetic will tell you whether you were a forecaster or a coin.

The Forecast Archive
The largest controlled dataset on expert judgment opens with a practice: writing probabilities down and scoring them. Over many years, Philip Tetlock collected probabilistic forecasts from experts—academics, government analysts, journalists—for the Expert Political Judgment project. According to Tetlock's book reporting the results, the experts' aggregated accuracy roughly matched simple extrapolation from base rates: the average specialist added little over simply asking how often this kind of event had happened before. Ideological "hedgehogs" running one big theory scored reliably worse than eclectic "foxes" stitching together many partial models.
The sharpest reversal came in IARPA's ACE tournament. According to Tetlock and Gardner's Superforecasting, the Good Judgment Project's superforecasters—the top performers among thousands of volunteers—posted season-average Brier scores near 0.14, beating the intelligence-community baseline by roughly 30% and prediction-market aggregates by double digits. Nothing in that winning profile resembles the hovering hedger. The top performers committed, frequently in the 0.7–0.9 range once named evidence justified it, and stayed revisable between updates. A probability pinned permanently at 0.5 has zero resolution—it scores identically whether its questions resolve 10% or 90% of the time—which is why Tetlock's tournament data shows perpetual hedgers finishing behind participants who simply quoted the base rate.
Calibration is trainable, not a personality trait. According to Mellers et al.'s PNAS report on experimental arms within the same tournament, volunteers assigned to probabilistic-training conditions and volunteers assigned to team-workshop conditions each lifted accuracy by roughly 10% or more over controls. Two independent interventions, one direction of effect: decomposing questions into reference classes and structuring group judgment both move Brier scores.
The 0.1–0.9 guardrail exists because untrained judgment fails in the opposite direction. According to Alpert and Raiffa's experiments, subjects instructed to give 98% confidence intervals captured the true value far less often than their stated confidence implied. People do not naturally hedge; they naturally overreach, converting thin evidence into near-certainty. Holding published probabilities inside the band caps the blast radius of that instinct, releasing the cap only when a specific causal mechanism—not vibes—justifies the extreme value.
Near-perfect calibration is demonstrably achievable. In the verification datasets analyzed by Allan Murphy and Edward Winkler, National Weather Service precipitation-probability forecasts plot nearly on the diagonal of the reliability curve—when the service publishes 30%, rain follows roughly 30% of the time. The mechanism is not talent but loop speed: forecasters issue scored probabilities daily, face verification almost immediately, and cannot delete misses from the record. Fast, repeated, scored feedback is what straightens a reliability curve.
The repair is fast, too. According to Douglas Hubbard's How to Measure Anything, after roughly half a day of calibration exercises with immediate feedback, professionals' 90%-confidence interval questions hit about 85–90% of the time—most of the hedgehog-to-fox gap closes in an afternoon. Before publishing your first probability of the quarter, run one such drill; then anchor each question in its reference-class base rate, adjust for named evidence, hold the band, and log every number for quarterly Brier scoring.
| Evidence | Sample and setting | Headline figure | What it licenses |
|---|---|---|---|
| Tetlock, Expert Political Judgment | Multi-year archive of probabilistic forecasts from academic, government, and journalistic experts | Accuracy roughly equaled base-rate extrapolation; hedgehogs reliably worse than foxes | Anchor first on the reference class |
| Tetlock & Gardner, Superforecasting | Top performers among thousands of Good Judgment Project volunteers | Season-average Brier near 0.14; roughly 30% over the intelligence-community baseline; double digits over prediction markets | Commit inside the band; update in measured steps |
| Mellers et al., PNAS | Probabilistic-training arm vs. control | Roughly 10% or more accuracy lift | Calibration is a trainable skill |
| Mellers et al., PNAS | Team-workshop arm vs. control | Roughly 10% or more accuracy lift | Structured aggregation compounds the gain |
| Alpert & Raiffa | Subjects giving 98% confidence intervals | True value captured well below the stated confidence level | Untrained judgment is extremist; hence the 0.9 cap |
| Murphy and Winkler verification datasets | National Weather Service precipitation probabilities | Reliability curve plots nearly on the diagonal | Daily scored feedback yields near-perfect calibration |
| Hubbard, How to Measure Anything | Half-day drill with immediate feedback | 90%-confidence questions hit about 85–90% | Run the drill before your first published forecast |

Three Archetypes, One Winner
The Perpetual Hedger's 2026 scorecard is finished before the year's first question resolves. Stating 0.4–0.6 on every question keeps the squared-miss penalty nearly constant whether an event lands or not, so his final average depends on question count alone—no skill can enter the calculation. His signature move, waiting for certainty before committing, delivers the failure mode exactly as designed: zero resolution, answers that carry no information about which outcomes occur, and a locked-in mediocre Brier score. In Tetlock's tournament archive, flat-profile forecasters were precisely the ones showing no resolution and losing to plain base-rate quoters. Hence the label, Aristotle misread: the doctrine of the mean locates virtue relative to a measurable midpoint between excess and deficiency, not at permanent indecision, and the belief that calibrated wisdom means hovering near 50/50 until certainty arrives gets the philosophy exactly backwards.
The Extremist Hedgehog loses by the mirror-image route. Narrative conviction drives stated probabilities of 0.85–0.99, and the bold call becomes the brand. The arithmetic is unforgiving: under the squared-error rule defined above, a 0.9 call that misses costs (1 − 0.9)² = 0.81—the same penalty as nine correct calls made at 0.7, which run (1 − 0.7)² = 0.09 apiece. Every surprise bleeds like that. The compensation is real: on genuine breaks—regime shifts, discontinuities, the events nobody else will touch—he wins spectacularly, posting near-zero penalties while everyone else clusters mid-scale. But genuine breaks are rare by definition, so across an ordinary year the bleed dominates the windfalls. Best-fit domain: rare regime shifts. Verdict: LOSE on average.
The Base-Rate Fox wins because he starts where frequencies live. He opens each question at the reference-class rate—whatever share of comparable past cases resolved yes—then states 0.55–0.85 as named evidence accumulates, revising in equal log-odds steps so identical evidence moves him identical distances in either direction. His weakness is specific and honest: slowness on true discontinuities, where stepwise updating trails the hedgehog's leap. But the broad run of resolvable 2026 questions is not made of discontinuities; it is made of base rates plus incremental news, precisely the terrain where anchoring-and-adjusting beats both paralysis and bravado. Verdict: WINNER.
When two archetypes post similar Brier scores, break the tie on resolution, not on feel. The classical decomposition of a proper score—reliability minus resolution plus irreducible uncertainty, introduced for binary forecasts by Allan Murphy—separates two qualities readers constantly conflate. Reliability asks whether stated probabilities match observed frequencies; it is largely a scoring artifact of how answers cluster into bins, and a forecaster who never moves can look beautifully reliable. Resolution asks whether the answers themselves vary with outcomes, and it is the component that improves with practice and sharper reference classes. Audit yourself at each quarterly scoring pass: split your resolved calls into those above and below your own median probability. If the high bucket did not resolve "yes" more often than the low bucket, your score is riding the artifact—you have collected reliability without acquiring the skill term.
One boundary condition keeps the comparison honest. These archetypes apply only to resolvable, time-stamped binary 2026 questions—will the fed funds target sit below 3.5% at year-end 2026, yes or no—where an eventual outcome settles every dispute. That constraint is structural: foundational treatments of forecast calibration begin with binary classification, in which the forecasted outcome takes exactly two values, true or false, 1 or 0. Open-ended value disputes—whether a policy is wise, whether a technology is culturally good—generate no resolution event, so no archetype can be scored there and the table below simply does not apply.
| Archetype | Stated range | Signature move | Failure mode | Best-fit domain | Verdict |
|---|---|---|---|---|---|
| Perpetual Hedger ("Aristotle misread") | 0.4–0.6 on every question | Waits for certainty | Zero resolution; mediocre score locked in before events resolve | None | LOSE |
| Extremist Hedgehog | 0.85–0.99, narrative conviction | The bold call | Bleeds 0.81-point penalties on every surprise; wins big only on genuine breaks | Rare regime shifts | LOSE on average |
| Base-Rate Fox | Opens at the reference-class frequency; 0.55–0.85 as evidence accumulates | Stepwise log-odds updates | Slowness on true discontinuities | Broad run of resolvable 2026 questions | WINNER |
To place yourself on the table: take one quarter of logged calls, run the median-split audit above, and compare your Brier average with what a pure reference-class anchor would have scored on the same questions—the larger gap tells you whether to repair your anchors or your updating.

What the Data Doesn't Tell You
Before you lean on any calibration result, ask who picked the questions. Both evidence bases behind the mean's win — the archive profiled above and the IARPA-sponsored tournaments — used question sets assembled by program staff, not sampled from everything you will actually be asked to predict. The slant runs toward geopolitics, macroeconomics, and public health, and most questions resolved well inside a year. The margin is real, but it was earned on their mix; whether it transfers to a sheet weighted toward technology adoption, corporate behavior, or your own specialty is exactly what the record leaves open.
The measurement has blind spots too. According to the published analyses from the IARPA-era research teams, much of the headline accuracy came from teams and pooled aggregates, not lone forecasters — the result certifies a workflow, not merely a temperament. And a quarterly Brier average over a handful of personal questions carries almost no signal; the score stabilizes only across many questions, which is why the archives needed enormous volumes before rankings meant anything. One person's season is noise wearing a spreadsheet.
Variance across cases is the quieter problem. Calibration is not a portable trait: the tournament records show individual forecasters whose accuracy swung sharply by genre — strong on political questions, mediocre on economic ones, or the reverse. A seasonal average buries that. Because the Brier score decomposes into reliability and resolution terms, a forecaster can post a respectable average while being systematically miscalibrated on an entire slice of their sheet; the average forgives what the slice reveals. The practical consequence: the anchor's quality tracks the reference class, not the forecaster. A base rate built on dozens of resolved analogues behaves nothing like one built on a handful.
So when does the rule itself break? In narrow, nameable places. First-of-a-kind events — no resolved analogues, no class — leave the base-rate half with nothing to grip; there the causal-mechanism citation carries the weight, and the data cannot audit whether your mechanism is physics or autobiography. Structural breaks stale-date the anchor; measured-step updating handles that, but only if each step's trigger is logged. And the premium above 0.9 is justified only when a specific causal mechanism is citable and the base rate already sits at the extreme — absent both, the band holds. Each is an edge case with a protocol, not a refutation.
None of this licenses the oldest dodge in forecasting: hovering at 0.5 until certainty arrives. Thin evidence argues for a coarser anchor and a louder uncertainty note in the log — never for abandoning the number. Perpetual even-odds has zero resolution and scores poorly whether you know something or nothing, as the archetype comparison above established. In every edge case below, the anchored, logged, revisable number beats the 0.5 default. The limitations change how honestly you describe your uncertainty, not whether you commit.
| Situation on your sheet | What the record supports | The move | Residual risk |
|---|---|---|---|
| Dense class: dozens of resolved analogues | Anchor is load-bearing | Publish the base rate; adjust only by named evidence | Misreading the analogy |
| Thin class: a handful of analogues | Anchor is coarse | Publish inside the band; widen the uncertainty note in the log | Small-sample noise |
| No class: first-of-a-kind event | Nothing to anchor | Cite the causal mechanism in the log; flag for review scrutiny | Mechanism as autobiography |
| Structural break: base rate predates the change | Anchor is stale | Update in measured steps, each tied to a logged trigger | Whipsawing on noise |
| Extreme base rate plus citable mechanism | The only licensed route past 0.9 | Document the mechanism before publishing the premium | Post-hoc rationalization |
Action for the current cycle: before the next quarterly scoring cut, tag every open 2026 entry with one of the five rows above. Anything you cannot tag is a malformed question — rewrite it or drop it — and is never a reason to park it at 0.5.

Also worth reading: Why waiting for certainty is the worst judgment call: Why waiting for certainty is · How to develop better judgment in a world full of distractions: How to develop better judgment · Why church attendance does not increase atheist acceptance because of the suppression effect: Why church attendance does not
Where the Mean Breaks
Every calibration curve behind the mean strategy was earned somewhere with fast feedback. The National Weather Service posts near-diagonal reliability because its forecasters issue thousands of scored probability-of-precipitation calls and get graded within a day. Your 2026 sheet has no such engine—a geopolitical or AI question usually resolves once. Tetlock's Expert Political Judgment fieldwork closed in a pre-smartphone, pre-LLM world, so the fox edge may compress wherever feedback loops run slow. Treat the mean as imported technology, not native law.
The second break is crowd imitation. Here is the mechanism the extremization literature exploits: pooled forecasts hug the middle because individual contributors are underconfident, so aggregators push averages outward. A solo forecaster who copies those extreme outputs applies the correction twice. According to Poly Syncer, crypto-price binaries on Polymarket are systematically underconfident at exactly the extreme probabilities where aggressive forecasters live, and public sportsbook lines on equivalent events run roughly triple Polymarket's 2.1-point calibration error. When the crowd itself needs correcting, its tails are a warning, not a template.
Third, thin tails. Below 0.10—your banking-crisis question, your forced-restructuring-of-a-frontier-AI-lab question—historical reference classes turn sparse, and post-2020 regime shifts structurally break whichever class you pick. Base-rate anchoring silently inherits that bias. The rot repeats up top: according to PredictEngine's February 28, 2026 analysis, events people rate 90% certain typically occur only 70–80% of the time. A broken class yields a broken anchor.
Fourth, the selection effect. Superforecasters were self-selected, heavily practiced volunteers answering IARPA-style questions—survivors of thousands of graded repetitions. Their roughly 0.14 Brier average is an upper bound, not a benchmark. A novice running the mean through 2026 should expect to land worse and improve quarter over quarter; the strategy sets a direction, not a floor.
Fifth, statistical power. One calendar year of 20–50 questions cannot separate a 0.19 fox from a 0.22 extremist—the confidence intervals overlap by construction. Small biases hide inside that noise: according to PredictEngine, a forecaster who prices 55% events at 70% systematically overpays for shares, and a single year's Brier average rarely exposes a gap that size. The metric itself bends too—Foster and Vohra demonstrated in 1998 that adversarial relabeling can drive measured calibration error to zero. 2026 is round one of a recurring discipline, not a verdict.
Last, the incentive wrinkle. Under quadratic loss, a forecaster uncertain about their own skill rationally shrinks estimates toward 0.5—some apparent hedging is optimal statistics, not cowardice. That is not a license to hover at 50/50 until certainty arrives. Rational shrinkage starts from a committed base-rate estimate and pulls it partway toward the middle in proportion to your demonstrated error; the perpetual-hedger pathology skips the estimate entirely. The cure is instrumentation: log every call, score quarterly, estimate your personal shrinkage factor, and tighten it as the record grows.
| Failure mode | Symptom on your 2026 sheet | Correction | ||||||||
| Slow feedback loops | One-shot geopolitical and AI questions; EPJ-era edge may compress | Weight questions with interim signals higher | ||||||||
| Crowd imitation | Extreme market prints copied onto a solo sheet | Aggregate sources once, apply one correction | ||||||||
Broken reference classes Frequently Asked QuestionsHow accurate were Polymarket prices compared to other public forecasts? Poly Syncer's May 17, 2026 audit of markets resolving between January 2024 and May 2026 found just 2.1 points of mean absolute calibration error across all probability buckets, sampled 24 hours before resolution — ahead of sportsbook lines and polling averages. When I feel 90% sure something will happen, how often does it actually happen? Events people rate 90% certain typically occur only about 70% to 80% of the time, per PredictEngine, and Convexly notes overconfidence clusters in the high-confidence bins of reliability diagrams. Isn't answering 50% on everything safe, since it keeps you perfectly calibrated? Answering 50% on every question delivers perfect calibration with zero discrimination — 'useless, but honest,' per EdgeMarket — so even a lazy forecaster who merely quotes the historical base rate, say 11.2%, outscores the fence-sitter. Should I extremize my own forecasts to sharpen them? No — extremization is a crowd correction, so individuals applying it simply import overconfidence; the winning Hedger instead bounds conviction so a genuine edge reads 55%, not 90%, while exploiting soft spots like systematically underconfident crypto-price binaries. What Brier scores did the best competitive forecasters actually achieve? In IARPA's ACE tournament, the Good Judgment Project's superforecasters posted season-average Brier scores near 0.14, beating the intelligence-community baseline by roughly 30% and prediction-market aggregates by double digits. Why does the scoring guide cap published probabilities inside the 0.1–0.9 band? Because untrained judgment fails toward overreach — Alpert and Raiffa's subjects instructed to give 98% confidence intervals captured the true value far less often than their stated confidence implied — so holding published probabilities inside the band caps the blast radius. Quick answers
Sources: arXiv, Reddit, arXiv, Reddit, Reddit Research Methodology & Editorial StandardsWe begin by defining the specific objectives the reader needs to accomplish. Primary product documentation and authoritative secondary sources are assembled into a verified research corpus; drafting occurs only after this foundation is in place. Every quantitative claim is subjected to dual-source verification. Any figure that cannot be independently corroborated is either qualified or omitted. Judgment Call Podcast Essays for people who make the callTechnology, philosophy, and society — long-form analysis for high-stakes judgment under uncertainty. Browse latest essays |