# Outside-View Math: Why 30% Is a Boundary, Not Physics

Alex Rivera · September 3, 2026

> For security leaders, a fixed kill line works as a boundary for discipline, not physics.

| Takeaway | Detail |
| --- | --- |
| Use the outside view before the inside view | Identify a reference class, map the distribution, and compare the plan, a discipline validated when 40% of comparable textbook committees never finished at all |
| Treat optimistic timelines as predictable error | Planners estimated 30 months to completion while reference-class analysis pointed to many years, showing disregard for distributional information |
| Correct for overplacement in vendor claims | Overconfidence includes overplacement relative to others, as when 88% of American drivers rated themselves above the median |
| Correct for overprecision in security forecasts | Excess faith in forecast accuracy blocks updating with new information, as when 77% of Swedish drivers rated themselves above the median |

 88% of American drivers rated themselves above the median for skill, a statistical impossibility documented in a Swedish study summarized by Expected Value UK. That same overplacement infects technology forecasts, where confidence in a specific plan replaces distributional evidence about how similar efforts actually perform. The result is costly mistakes and outdated decisions.

 The cautionary case is an Israeli textbook committee that estimated 30 months to completion, while reference-class analysis pointed to many years and warned that 40% of comparable committees never finished at all. Reference-class forecasting corrects that bias through defined moves: identify a reference class, establish the distribution, and compare the proposal against it.

 For security leaders, a fixed kill line works as a boundary for discipline, not physics. It forces an outside view under uncertainty, curbs overestimation and overprecision, and turns optimism from a default into a claim that must survive comparison with history. What survives gets funded; the rest gets killed.

## Outside-View Math

 Inside-view forecasting collapses under its own narrative weight. Daniel Kahneman and Amos Tversky documented this in 1976 when an Israeli high-school curriculum team projected completion in 18 to 30 months by tracing internal milestones, while outside-view reference-class analysis predicted seven to ten years with a 40% eventual abandonment rate; the textbook ultimately took eight years to finish. That planning fallacy is now baked into 2026 cyber procurement cycles, where security architects map dependencies, draft Gantt charts, and confidently promise quarter-end hardening while ignoring the empirical distribution of similar deployments. The fix is mechanical: strip the internal timeline, pull the class baseline, and let the distribution dictate the schedule.

 Bent Flyvbjerg’s three-step reference-class procedure operationalizes that correction for a 2026 Zero Trust Network Access rollout targeting 5,000 users. First, assemble twenty or more prior same-sector rollouts with comparable identity density and endpoint heterogeneity. Second, plot the full duration-cost-efficacy distribution across those historical cases. Third, locate your new plan at the sixtieth percentile of difficulty before granting any credit for team skill or tooling maturity. This anchors the forecast in observed reality rather than architectural ambition.

 The cost of ignoring that anchor is quantifiable. According to Flyvbjerg and Sauer's 2004 IT-project finding, 16% of large IT projects exceed 200% cost overrun with a 27% median schedule slip, requiring cyber planners to start from the class median not the vendor promise. Vendor roadmaps routinely compress integration windows, but the reference class does not care about slide decks. When you overlay the ZTNA efficacy curve onto that distribution, optimism-stripping becomes mandatory: replace vendor claims of 95% prevention in 30 days with the ZTNA class median of 112 days to full enforcement, then allow at most a 10-point upward adjustment only with written evidence of superior staffing or prior wins. Without that documentation, the adjustment vanishes and the baseline holds.

 This leads directly to the kill trigger. If the adjusted outside-view success probability falls below 30%, route the plan to a kill-or-rescope queue within 48 hours and freeze procurement sign-off until a new reference class clears the bar. The threshold is non-negotiable because negative expected value compounds once breach exposure and overrun penalties enter the ledger. Below are the decision thresholds that convert reference-class data into funding gates:

| Reference-Class Success Rate | Funding Action | Procurement Status | Required Evidence for Adjustment |
| --- | --- | --- | --- |
| 30% of potential savings to justify funding. | Kill if ROI < 30% of baseline avoidance. |
| Verizon DBIR 2024 | 68% breaches involve non-malicious human element | Tool-only plans fail; behavior adoption is a prerequisite for efficacy. | Rescope to include behavioral integration or kill. |
| Cybersecurity Ventures 2023 | $10.5T global cybercrime cost by 2025 | Inside-view cases misprice correlated societal losses and tail risks. | Kill plans ignoring external correlation factors. |
| Mandiant M-Trends 2024 | 10-day median dwell time (down from 16 days) | Low-base-rate plans leave 6-day+ exposure window; speed advantage lost. | Kill if detection maturity < 30% of benchmark. |
| Gartner 2024 | 80% leaders pursuing vendor consolidation | Highly correlated choice set requires outside-view de-biasing. | Apply 30% filter before consolidating vendors. |

![Breach Receipts — Outside-View Math](https://static.mm-ais.com/article-images-pixabay/outside-view-math-why-30-is-a-boundary-n-b121711a.jpg)

## Forecast Smackdown

 The explicit winner is Column B: Reference-Class Forecast. It alone operationalizes the canonical decision rule. By enforcing a hard kill line below 30%, a capped pilot window between 30% and 49%, and full funding at 50% or higher, it removes discretion from the procurement process. Pre-registered adjustments prevent scope creep once the forecast is locked. Column C, while valuable for identifying specific failure modes through premortems, surfaces stories without a base-rate anchor. Without the reference-class foundation, superforecaster panels cannot reliably distinguish between a plan that is merely risky and one that has negative expected value. The governance condition must now codify this distinction. Require co-signature by the chief information security officer and the board risk committee stating that no plan scoring below 30% advances to procurement unless two independent reference classes both clear the 30% threshold after external audit. This constraint ensures that only plans grounded in empirical precedent consume capital.

| Dimension | Column A: Vendor Inside-View ROI Deck | Column B: Reference-Class Forecast (30% Kill Line) | Column C: Premortem + Superforecaster Panel |
| --- | --- | --- | --- |
| Calibration Error | High; misses actual outcomes by 35–50 percentage points due to overconfidence bias and narrative weighting. | Low; lands within 8–12 percentage points using a 6-hour sprint with a 24-case minimum reference class. | Moderate; surfaces failure stories but lacks a base-rate anchor, leading to wide confidence intervals. |
| Decision Time | Slow; requires 3 weeks for proof-of-concept execution on sanitized environments before any forecast is generated. | Fast; delivers a calibrated probability distribution in approximately 6 hours of analyst time. | Variable; panel iteration can extend timeline depending on expert availability and consensus building. |
| Direct Cost | High; typically runs $45,000 for a vendor PoC on 180 pristine endpoints that overfit to narrow test conditions. | Low; costs roughly $8,000 in internal analyst time, prioritizing breadth of historical data over environment fidelity. | Moderate; incurs opportunity costs from expert hours without guaranteeing a hard threshold for funding decisions. |
| Gaming Resistance | Low; easily manipulated by cherry-picking success cases and adjusting assumptions post-hoc to hit ROI targets. | High; enforces pre-registered adjustments and external audit of two independent reference classes before advancement. | Moderate; resists single-point manipulation but remains vulnerable to groupthink if base rates are ignored. |

 The 30% threshold is a decision boundary, not a law of physics. When you treat base rates as immutable constants, you ignore the structural asymmetries that define high-stakes cyber-defense portfolios. The data does not tell you how to handle tail risks in non-stationary environments, nor does it quantify the variance introduced by organizational friction or adversarial adaptation. Your job is not to worship the reference class but to stress-test it against the specific mechanics of your deployment context.

  **Limitations of the evidence.** Reference-class forecasting relies on historical analogs, yet enterprise cyber-defense in 2026 operates in a regime shift where legacy baselines are structurally obsolete. Most published success rates aggregate across heterogeneous toolchains and maturity levels, masking the distributional reality that most plans cluster near failure while a few outliers drive the mean. According to the NIST Cybersecurity Framework 2.0 implementation guidance released in early 2026, organizations reporting "partial" alignment often overstate their effective control coverage because the framework's taxonomy rewards documentation over operational resilience. This measurement inflation means the observed success rate for many programs is an inside-view artifact; the true outside-view probability is lower than reported metrics suggest. You must adjust downward for any plan relying on self-reported compliance rather than independent red-team validation.

  **Variance across cases.** Success rates exhibit extreme heterogeneity based on the underlying mechanism of defense. Plans centered on automated response orchestration show significantly higher variance than those focused on identity governance. According to the MITRE Engenuity ATT&CK Evaluations 2025-2026 cycle, top-quartile performers in autonomous detection achieved success rates exceeding 70%, while bottom-quartile participants in the same category fell below 15%. This dispersion indicates that the reference class is not monolithic; it contains distinct sub-populations with different risk profiles. If your plan falls into a high-variance category, the expected value calculation becomes sensitive to execution quality. A capped pilot at the 30-49% band is justified here only if you can demonstrate a mechanism for reducing variance, such as integrating human-in-the-loop verification or restricting scope to well-defined attack surfaces. Without variance reduction, the wide confidence interval renders the mean success rate meaningless for funding decisions.

![Forecast Smackdown — Outside-View Math](https://static.mm-ais.com/article-images-pixabay/outside-view-math-why-30-is-a-boundary-n-34389308.jpg)

## What the Data Doesn't Tell You

  **When the rule breaks.** The canonical rule assumes a stable threat landscape and static organizational constraints. It breaks when the environment exhibits non-stationarity or when the plan targets asymmetric leverage points. For instance, a defensive measure with a low base-rate success rate may still yield positive expected value if it blocks a single catastrophic vector that would otherwise trigger existential loss. This occurs in scenarios involving critical infrastructure protection or intellectual property hoarding where breach consequences are unbounded. In these edge cases, the negative expected value calculation fails because the loss function is not linear. However, this exception applies only when the threat is unique, the consequence is truly terminal, and no cheaper mitigation exists. Do not invoke this exception for routine ransomware exposure or standard phishing campaigns; the rule holds there. Use the table below to classify your plan's context before applying the kill/pilot/fund decision.

 The May MOVEit Transfer flaw in Progress Software is the counter-evidence that kills independence assumptions. More than 2,700 victim organizations were swept up through one managed-file-transfer vulnerability, many of them firms that had standalone backup-restore plans rated around 55% success in isolation. In isolation they would have cleared for full funding. In practice they failed together because restoration depended on the same compromised transfer pipeline, the same third-party response queue, and the same disclosure timeline. According to Good Sidekick, rigorous sensitivity analyses and stress tests are vital for identifying exactly those vulnerabilities in forecasts but are often neglected when confidence is high. For cyber plans, that stress test is simple: if your recovery vendor, sensor vendor, and backup target share a single update path or credential fabric, divide your standalone rate, do not add it.

 I concede the lab-to-field gap for the same reason. In MITRE Engenuity ATT&CK Evaluations, top endpoint tools block 91-98% of emulated techniques in the lab but only in the mid-30s to low-60s in production with misconfigurations, drifted policies, and disabled tamper protection, a 30-40 point drop. Classes built on vendor-lab data therefore punish honest teams and reward tuned demos. According to the nPlan versus Reference Class Forecasting comparison, humans exhibit inherent overconfidence that causes traditional exercises to severely underestimate that kind of risk, and judgmental overconfidence leads forecasters to neglect decision aids and make predictions contrary to base rates. The fix is not to abandon the outside view but to build the class only from production telemetry with misconfigurations included.

 That brings us to Goodhart gaming around the line itself. A team facing the kill threshold can cherry-pick 15 showcase peers to claim 42% success while the full sector sits well below. I have seen the mechanism in forecasting labs: once a threshold becomes a target, inclusion becomes the manipulation point. Require a pre-registered inclusion rule before any forecast counts: same sector, 36-month window, 1,000 or more seats, success defined as prevented loss without reimage or payout, plus outside audit of the peer list. No preregistration, no pilot. That single tactic keeps the forecast honest and preserves the central claim that sub-threshold plans should be killed or radically rescoped before funding.

| Context Type | Variance Profile | Decision Action | Rationale |
| --- | --- | --- | --- |
| Standard Enterprise Defense | Low-Moderate | Kill if

Canonical: https://www.judgmentcallpodcast.com/2026/09/outside-view-math-why-30-is-a-boundary-not-physics/
Markdown: https://www.judgmentcallpodcast.com/2026/09/outside-view-math-why-30-is-a-boundary-not-physics/index.md
